Station 07Open the adventure map
00Set up your toolsSetup01The starting lineSetup02Processes and mailboxesR1An Elixir data pipelineOptional reviewR2Read ErlangOptional review03Messages and timeouts04OTP rules for messages05A supervision tree that restarts06Put a limit on concurrency07When a BEAM node disappearsX1Agree on shared BEAM termsBilingual extensionX2Two-language partnersBilingual extension08A reliable job runner
Home/BEAM mainline/Station 07
07
NetworksNetwork challengeBEAMOTP

When a BEAM node disappears

Connecting nodes is only the beginning. Plan for disconnections, late replies, and stale data too.

4checkpoints
About 6 hours · Try 4 sessionssplit it into sessions
QUESTION · ONE PROBLEM FOR THIS STATION

If a remote worker disconnects midway through a request, what result does the caller get, and how does the system mark stale data, retry, and avoid duplicate execution?

Start at the scene

The node was gone, but the dashboard stayed green

A remote node disconnects midway through a request, yet its last metric still appears healthy. If the caller retries, the original task may also have completed and the external action can happen twice.

Evidence you can observe
  • A timeline of nodeup, nodedown, and request references
  • The last sample timestamp and its stale marker
  • Retry and idempotency records plus registration logs after reconnect
Why this station matters

The problem it solves

Distributed Erlang makes remote sends look natural, but a network can still delay or disconnect. Node names and cookies only help connections. We must still design consistency, delivery, capacity, and security.

After this station

You will be able to

  • Start two named BEAM nodes and let them find each other and exchange messages

  • Explain nodedown and separate a remote PID from a registered name that works on only one node

  • Use contextual logs, Telemetry metrics, and Observer to track mailboxes and delay

Before you start
  • Know how processes communicate, supervision trees recover, and backpressure reports overload
  • Be able to start a basic OTP Application. You can continue even without making a release yet
One job, two ways to write it

See what it does before how it is written

Connections and disconnections become messages. Agree in advance how to degrade or reject work after nodedown.

Current code
Elixir
# Ask the current process to receive node connection events
:net_kernel.monitor_nodes(true)

# Node events arrive like ordinary BEAM messages
receive do
  {:nodeup, node} ->
    # Record the node that just connected
    Logger.info("node connected", node: node)

  {:nodedown, node} ->
    # The application chooses whether to degrade or retry
    Logger.warning("node disconnected", node: node)
end
DESIGN DECISION

Why this shape?

Why design a disconnect protocol if nodes can already send messages?

Choice

A convenient send syntax helps only after connection. Networks still delay and partition, and old data can look current.

Cost and boundary

Retries improve the chance of success but can repeat effects, so stale markers, idempotency, and retry limits travel together.

A FAMILIAR POINT OF VIEW

Coming from Java, Python, or JavaScript

Java

Familiar starting point
RPC, brokers, or clustering frameworks.
What BEAM changes
Sending to a remote PID looks close to local send, but its failure semantics are still network semantics.
False friend
A convenient call form does not provide exactly-once delivery.

Python

Familiar starting point
HTTP/RPC or task queues such as Celery.
What BEAM changes
Node events also arrive as BEAM messages the application can handle.
False friend
A nodedown event does not explain why the connection failed.

JavaScript

Familiar starting point
fetch, WebSocket, or queue consumers.
What BEAM changes
The same VM model reaches across nodes, but boundaries still need timeouts and idempotency.
False friend
Short syntax cannot erase a network partition.
LAB
Hands on

Take one node offline

About 15–25 minutes

Start two local nodes, connect them, then stop one. Watch the event, remote processes, and unfinished requests.

  1. 01

    Start a@127.0.0.1 and b@127.0.0.1 with the same cookie

  2. 02

    From A, ping B and turn on monitor_nodes

  3. 03

    Stop B, record A's event, and see how unfinished requests end

Copy into the terminal and press Enter
# Try node b; success returns :pong and failure returns :pang
Node.ping(:"b@127.0.0.1")
What you should see
  • After a successful connection, Node.ping/1 returns :pong
  • After B stops, A receives nodedown
  • BEAM reports the disconnection, but the application still decides whether a request fails, retries later, or uses another path
Break it on purpose

Give the two nodes different cookies. Confirm that the failure happens during connection rather than in a business handler.

What this shows

Node connections and disconnections can be observed and turned into events an application can handle.

What this does not show yet

Stopping a local node cannot simulate packet loss, long delay, half-open connections, or a network split.

Meet the words

Key ideas in this code

01

node

A BEAM instance taking part in distributed communication, usually named name@host. A remote PID also contains its node identity.

02

cookie

A shared secret used when nodes connect. It is not a complete security plan and cannot replace encryption, isolation, and access control.

03

Telemetry

An event-measurement tool often used by Elixir. Business code emits events; handlers count or export them.

Name the shape in the code

Design patterns used here

01

Circuit Breaker

Stop routing to an unavailable node and probe to decide when it can recover.

02

Lease

Expire health and registration facts over time instead of treating old data as current.

03

Idempotent Consumer

Keep a retry or duplicate delivery from repeating an external effect.

Think it through

Distributed Erlang makes sending to a remote PID look like a local send. What does it provide automatically?

Your turn

Draw a node-status map

Report scheduler use, memory, and important mailboxes regularly. Keep the last value after a disconnect and mark it stale.

Hint 1Take the first step

Add a sample time to every metric. A number without its time can mislead.

Hint 2Make it a little smaller

Keep the last real data after disconnection and mark it clearly as stale.

Hint 3You are close

Watch the trend of a growing mailbox, not only one number at one instant.

Hint 4Work backward from the finish

Choose one success signal and write the smallest test for it. If the computer cannot show the result, rewrite the signal as something you can truly observe.

Ready to move on when
  • A disconnected node does not appear as “every metric is a healthy zero”

  • Logs show which node, process, and request/reference each event belongs to

  • Reporting resumes after reconnection without registering the same service twice

Take these with you

Remember three things

  1. 1

    A remote send looks like a local send for convenience. The network has not disappeared.

  2. 2

    Disconnections, timeouts, and stale data belong in the agreement before trouble begins.

  3. 3

    Besides “the process is alive,” mailbox size, delay, restart count, and rejection rate say more about real health.

Read a little more

Visit the original sources

Distributed ErlangElixir LoggerTelemetry
Station completeYou ran the experiment and thought through the answer. Save this station.