Station 07Open the adventure map
When a BEAM node disappears
Connecting nodes is only the beginning. Plan for disconnections, late replies, and stale data too.
If a remote worker disconnects midway through a request, what result does the caller get, and how does the system mark stale data, retry, and avoid duplicate execution?
The node was gone, but the dashboard stayed green
A remote node disconnects midway through a request, yet its last metric still appears healthy. If the caller retries, the original task may also have completed and the external action can happen twice.
- A timeline of nodeup, nodedown, and request references
- The last sample timestamp and its stale marker
- Retry and idempotency records plus registration logs after reconnect
The problem it solves
Distributed Erlang makes remote sends look natural, but a network can still delay or disconnect. Node names and cookies only help connections. We must still design consistency, delivery, capacity, and security.
You will be able to
Start two named BEAM nodes and let them find each other and exchange messages
Explain
nodedownand separate a remote PID from a registered name that works on only one nodeUse contextual logs, Telemetry metrics, and Observer to track mailboxes and delay
- Know how processes communicate, supervision trees recover, and backpressure reports overload
- Be able to start a basic OTP Application. You can continue even without making a release yet
See what it does before how it is written
Connections and disconnections become messages. Agree in advance how to degrade or reject work after nodedown.
# Ask the current process to receive node connection events
:net_kernel.monitor_nodes(true)
# Node events arrive like ordinary BEAM messages
receive do
{:nodeup, node} ->
# Record the node that just connected
Logger.info("node connected", node: node)
{:nodedown, node} ->
# The application chooses whether to degrade or retry
Logger.warning("node disconnected", node: node)
endWhy this shape?
Why design a disconnect protocol if nodes can already send messages?
A convenient send syntax helps only after connection. Networks still delay and partition, and old data can look current.
Retries improve the chance of success but can repeat effects, so stale markers, idempotency, and retry limits travel together.
Coming from Java, Python, or JavaScript
Java
- Familiar starting point
- RPC, brokers, or clustering frameworks.
- What BEAM changes
- Sending to a remote PID looks close to local send, but its failure semantics are still network semantics.
- False friend
- A convenient call form does not provide exactly-once delivery.
Python
- Familiar starting point
- HTTP/RPC or task queues such as Celery.
- What BEAM changes
- Node events also arrive as BEAM messages the application can handle.
- False friend
- A
nodedownevent does not explain why the connection failed.
JavaScript
- Familiar starting point
- fetch, WebSocket, or queue consumers.
- What BEAM changes
- The same VM model reaches across nodes, but boundaries still need timeouts and idempotency.
- False friend
- Short syntax cannot erase a network partition.
Take one node offline
Start two local nodes, connect them, then stop one. Watch the event, remote processes, and unfinished requests.
- 01
Start
a@127.0.0.1andb@127.0.0.1with the same cookie - 02
From A,
pingB and turn onmonitor_nodes - 03
Stop B, record A's event, and see how unfinished requests end
# Try node b; success returns :pong and failure returns :pang
Node.ping(:"b@127.0.0.1")- After a successful connection,
Node.ping/1returns:pong - After B stops, A receives
nodedown - BEAM reports the disconnection, but the application still decides whether a request fails, retries later, or uses another path
Give the two nodes different cookies. Confirm that the failure happens during connection rather than in a business handler.
Node connections and disconnections can be observed and turned into events an application can handle.
Stopping a local node cannot simulate packet loss, long delay, half-open connections, or a network split.
Key ideas in this code
node
A BEAM instance taking part in distributed communication, usually named name@host. A remote PID also contains its node identity.
cookie
A shared secret used when nodes connect. It is not a complete security plan and cannot replace encryption, isolation, and access control.
Telemetry
An event-measurement tool often used by Elixir. Business code emits events; handlers count or export them.
Design patterns used here
Circuit Breaker
Stop routing to an unavailable node and probe to decide when it can recover.
Lease
Expire health and registration facts over time instead of treating old data as current.
Idempotent Consumer
Keep a retry or duplicate delivery from repeating an external effect.
Distributed Erlang makes sending to a remote PID look like a local send. What does it provide automatically?
Draw a node-status map
Report scheduler use, memory, and important mailboxes regularly. Keep the last value after a disconnect and mark it stale.
Hint 1Take the first step
Add a sample time to every metric. A number without its time can mislead.
Hint 2Make it a little smaller
Keep the last real data after disconnection and mark it clearly as stale.
Hint 3You are close
Watch the trend of a growing mailbox, not only one number at one instant.
Hint 4Work backward from the finish
Choose one success signal and write the smallest test for it. If the computer cannot show the result, rewrite the signal as something you can truly observe.
A disconnected node does not appear as “every metric is a healthy zero”
Logs show which node, process, and request/reference each event belongs to
Reporting resumes after reconnection without registering the same service twice
Remember three things
- 1
A remote send looks like a local send for convenience. The network has not disappeared.
- 2
Disconnections, timeouts, and stale data belong in the agreement before trouble begins.
- 3
Besides “the process is alive,” mailbox size, delay, restart count, and rejection rate say more about real health.