Station 05Open the adventure map
A supervision tree that restarts
First learn who depends on whom. Then decide which partners restart when one falls.
When Registry, a worker supervisor, and a dispatcher depend on one another, which restart strategy recovers enough without restarting too much?
One independent child exited, and the whole subsystem restarted
A supervisor puts independent children under one_for_all. When one worker exits, Registry and the dispatcher restart too; with the wrong order, a dependent process instead keeps stale state.
- The child dependency graph, start order, and supervision-tree structure
- PIDs before and after each injected exit
- Restart intensity plus state and external-effect checks after recovery
The problem it solves
“Let it crash” does not mean ignoring errors. A process that cannot continue lets a supervisor restart it. Expected cases such as a wrong password or low stock should still return normal results.
You will be able to
Draw a supervision tree from the questions “who depends on whom?” and “who should recover together?”
Compare
one_for_one,one_for_all, andrest_for_onein your own wordsChoose permanent, transient, or temporary from the expected response to normal and abnormal exits
- Write a simple GenServer or gen_server
- Know that links carry exit signals and that a process can end normally or abnormally
See what it does before how it is written
rest_for_one uses child order to express dependencies. A failure near the front restarts the dependents after it.
# Child order changes the recovery range of rest_for_one
children = [
# Registry starts first because later components use it
{Registry, keys: :unique, name: Jobs.Registry},
# The dynamic supervisor watches temporary workers
{DynamicSupervisor,
strategy: :one_for_one,
name: Jobs.Workers},
Jobs.Dispatcher
]
# Start the whole group under one root supervisor
Supervisor.start_link(
children,
strategy: :rest_for_one,
name: Jobs.Supervisor
)Why this shape?
Why restart under a supervisor instead of catching errors everywhere?
When a local process cannot continue, a supervisor can restore a known initial state according to real dependencies.
Expected business errors still return normally. Restarting cannot restore lost data or undo an external side effect.
Coming from Java, Python, or JavaScript
Java
- Familiar starting point
- Catching exceptions, container restarts, or replacing pool workers.
- What BEAM changes
- A supervision tree keeps process dependencies and local recovery inside the application.
- False friend
- Expected outcomes still return as data; not every error should crash.
Python
- Familiar starting point
- try/except, task groups, and external process managers.
- What BEAM changes
- Links and exit signals make supervision a runtime protocol.
- False friend
- A restart does not restore lost in-memory state.
JavaScript
- Familiar starting point
- catch, Worker exits, and process managers.
- What BEAM changes
- A supervisor chooses a recovery boundary from child specs and strategy.
- False friend
- Restarting cannot retract an HTTP request already sent.
Make each child exit
Record three PIDs. Crash the dispatcher and registry one at a time, then compare which PIDs change.
- 01
Start the supervision tree and record every child PID with
which_children - 02
Make the last child exit abnormally, then inspect the PIDs again
- 03
Make the first child exit abnormally and compare how many PIDs change this time
# List each child name, PID, type, and module
Supervisor.which_children(Jobs.Supervisor)- When the last child fails, usually only that child restarts
- When the first child fails,
rest_for_onerestarts it and every dependent after it
Change the strategy to one_for_all, then crash the independent last child. Watch for unnecessary restarts.
The restart strategy and child order together decide which processes start again.
A new PID shows only that a process restarted. Business state, disk data, and outside actions still need checking.
Key ideas in this code
Failure domain
The area affected by a failure. Some components recover together while others stay isolated. A supervision tree expresses that choice in its shape.
restart intensity
The largest number of restarts allowed during a time window. If the limit is passed, the supervisor exits and hands the problem to its parent.
Application
An OTP component that can start, stop, receive configuration, and declare dependencies. An application callback usually starts the root supervisor.
Design patterns used here
Supervision Tree
Encode process ownership, dependencies, and recovery boundaries as a hierarchy.
Bulkhead
Split independent work into separate failure domains so one failure does not restart everything.
Restart Strategy
Choose one_for_one, rest_for_one, or one_for_all from the dependency relationship.
Which relationship fits rest_for_one best?
Draw a supervision tree
Draw a supervision tree for a task queue. Mark every child's restart type and explain the recovery range.
Hint 1Take the first step
Find who owns long-lived state first, then draw its dependencies on other processes.
Hint 2Make it a little smaller
Short tasks and long-lived infrastructure usually should not use identical child specs.
Hint 3You are close
Record who receives the problem after the restart-intensity limit is passed.
Hint 4Work backward from the finish
Choose one success signal and write the smallest test for it. If the computer cannot show the result, rewrite the signal as something you can truly observe.
Every parent-child relationship has a recovery reason
Restart types match expectations for normal and abnormal exits
The diagram names at least one business error that should not be handled by crashing or automatic retry
Remember three things
- 1
A supervision tree first answers: who depends on whom, and where should a failure stop spreading?
- 2
“Let it crash” does not ignore expected errors. It hands a local problem that cannot continue to a supervisor.
- 3
Restarting a process is only one part of recovery. Data, outside actions, and idempotency still need careful design.