Station 05Open the adventure map
00Set up your toolsSetup01The starting lineSetup02Processes and mailboxesR1An Elixir data pipelineOptional reviewR2Read ErlangOptional review03Messages and timeouts04OTP rules for messages05A supervision tree that restarts06Put a limit on concurrency07When a BEAM node disappearsX1Agree on shared BEAM termsBilingual extensionX2Two-language partnersBilingual extension08A reliable job runner
Home/BEAM mainline/Station 05
05
OTPIntermediate explorationElixirErlangOTP

A supervision tree that restarts

First learn who depends on whom. Then decide which partners restart when one falls.

4checkpoints
About 6 hours · Try 4 sessionssplit it into sessions
QUESTION · ONE PROBLEM FOR THIS STATION

When Registry, a worker supervisor, and a dispatcher depend on one another, which restart strategy recovers enough without restarting too much?

Start at the scene

One independent child exited, and the whole subsystem restarted

A supervisor puts independent children under one_for_all. When one worker exits, Registry and the dispatcher restart too; with the wrong order, a dependent process instead keeps stale state.

Evidence you can observe
  • The child dependency graph, start order, and supervision-tree structure
  • PIDs before and after each injected exit
  • Restart intensity plus state and external-effect checks after recovery
Why this station matters

The problem it solves

“Let it crash” does not mean ignoring errors. A process that cannot continue lets a supervisor restart it. Expected cases such as a wrong password or low stock should still return normal results.

After this station

You will be able to

  • Draw a supervision tree from the questions “who depends on whom?” and “who should recover together?”

  • Compare one_for_one, one_for_all, and rest_for_one in your own words

  • Choose permanent, transient, or temporary from the expected response to normal and abnormal exits

Before you start
  • Write a simple GenServer or gen_server
  • Know that links carry exit signals and that a process can end normally or abnormally
One job, two ways to write it

See what it does before how it is written

rest_for_one uses child order to express dependencies. A failure near the front restarts the dependents after it.

Current code
Elixir
# Child order changes the recovery range of rest_for_one
children = [
  # Registry starts first because later components use it
  {Registry, keys: :unique, name: Jobs.Registry},
  # The dynamic supervisor watches temporary workers
  {DynamicSupervisor,
   strategy: :one_for_one,
   name: Jobs.Workers},
  Jobs.Dispatcher
]

# Start the whole group under one root supervisor
Supervisor.start_link(
  children,
  strategy: :rest_for_one,
  name: Jobs.Supervisor
)
DESIGN DECISION

Why this shape?

Why restart under a supervisor instead of catching errors everywhere?

Choice

When a local process cannot continue, a supervisor can restore a known initial state according to real dependencies.

Cost and boundary

Expected business errors still return normally. Restarting cannot restore lost data or undo an external side effect.

A FAMILIAR POINT OF VIEW

Coming from Java, Python, or JavaScript

Java

Familiar starting point
Catching exceptions, container restarts, or replacing pool workers.
What BEAM changes
A supervision tree keeps process dependencies and local recovery inside the application.
False friend
Expected outcomes still return as data; not every error should crash.

Python

Familiar starting point
try/except, task groups, and external process managers.
What BEAM changes
Links and exit signals make supervision a runtime protocol.
False friend
A restart does not restore lost in-memory state.

JavaScript

Familiar starting point
catch, Worker exits, and process managers.
What BEAM changes
A supervisor chooses a recovery boundary from child specs and strategy.
False friend
Restarting cannot retract an HTTP request already sent.
LAB
Hands on

Make each child exit

About 15–25 minutes

Record three PIDs. Crash the dispatcher and registry one at a time, then compare which PIDs change.

  1. 01

    Start the supervision tree and record every child PID with which_children

  2. 02

    Make the last child exit abnormally, then inspect the PIDs again

  3. 03

    Make the first child exit abnormally and compare how many PIDs change this time

Copy into the terminal and press Enter
# List each child name, PID, type, and module
Supervisor.which_children(Jobs.Supervisor)
What you should see
  • When the last child fails, usually only that child restarts
  • When the first child fails, rest_for_one restarts it and every dependent after it
Break it on purpose

Change the strategy to one_for_all, then crash the independent last child. Watch for unnecessary restarts.

What this shows

The restart strategy and child order together decide which processes start again.

What this does not show yet

A new PID shows only that a process restarted. Business state, disk data, and outside actions still need checking.

Meet the words

Key ideas in this code

01

Failure domain

The area affected by a failure. Some components recover together while others stay isolated. A supervision tree expresses that choice in its shape.

02

restart intensity

The largest number of restarts allowed during a time window. If the limit is passed, the supervisor exits and hands the problem to its parent.

03

Application

An OTP component that can start, stop, receive configuration, and declare dependencies. An application callback usually starts the root supervisor.

Name the shape in the code

Design patterns used here

01

Supervision Tree

Encode process ownership, dependencies, and recovery boundaries as a hierarchy.

02

Bulkhead

Split independent work into separate failure domains so one failure does not restart everything.

03

Restart Strategy

Choose one_for_one, rest_for_one, or one_for_all from the dependency relationship.

Think it through

Which relationship fits rest_for_one best?

Your turn

Draw a supervision tree

Draw a supervision tree for a task queue. Mark every child's restart type and explain the recovery range.

Hint 1Take the first step

Find who owns long-lived state first, then draw its dependencies on other processes.

Hint 2Make it a little smaller

Short tasks and long-lived infrastructure usually should not use identical child specs.

Hint 3You are close

Record who receives the problem after the restart-intensity limit is passed.

Hint 4Work backward from the finish

Choose one success signal and write the smallest test for it. If the computer cannot show the result, rewrite the signal as something you can truly observe.

Ready to move on when
  • Every parent-child relationship has a recovery reason

  • Restart types match expectations for normal and abnormal exits

  • The diagram names at least one business error that should not be handled by crashing or automatic retry

Take these with you

Remember three things

  1. 1

    A supervision tree first answers: who depends on whom, and where should a failure stop spreading?

  2. 2

    “Let it crash” does not ignore expected errors. It hands a local problem that cannot continue to a supervisor.

  3. 3

    Restarting a process is only one part of recovery. Data, outside actions, and idempotency still need careful design.

Read a little more

Visit the original sources

OTP Design PrinciplesElixir Supervisor
Station completeYou ran the experiment and thought through the answer. Save this station.