Station 08Open the adventure map
00Set up your toolsSetup01The starting lineSetup02Processes and mailboxesR1An Elixir data pipelineOptional reviewR2Read ErlangOptional review03Messages and timeouts04OTP rules for messages05A supervision tree that restarts06Put a limit on concurrency07When a BEAM node disappearsX1Agree on shared BEAM termsBilingual extensionX2Two-language partnersBilingual extension08A reliable job runner
Home/BEAM mainline/Station 08
08
ProjectCapstone projectElixirErlangOTP

A reliable job runner

Protect capacity and failure behavior in one language first. Add a second-language boundary as an extension.

4checkpoints
About 10 hours · Try 6 sessionssplit it into sessions
QUESTION · ONE PROBLEM FOR THIS STATION

How do we keep capacity and retries bounded through queue-full, worker exit, timeout, and node loss—and state honestly what is still missing across a full VM restart?

Start at the scene

One persistently failing job caused a retry storm

A failing job has no retry limit and no idempotency key. It returns to the queue again and again, capacity counters drift, and an external action may happen more than once.

Evidence you can observe
  • Queued, running, and retry counts before and after each fault
  • The idempotency key and a record proving the external effect happened once
  • Reject, recover, and give-up metrics plus injected-fault tests
Why this station matters

The problem it solves

The final project joins the earlier tools. We will fill the queue, fail a task, delay a message, and disconnect a node, then check state, retries, recovery, and rejection rules.

After this station

You will be able to

  • Write down rules that must always hold before arranging processes, queues, and supervision trees

  • Ship a one-language core first, then decide whether an Erlang/Elixir adapter adds real value

  • Create small failures on purpose and check how the system recovers, rejects work, or leaves clues

Before you start
  • Finish the preceding BEAM mainline and keep the queue and supervision exercises
  • The core needs one language; the bilingual extension also needs the shared-terms and interoperability stations
One job, two ways to write it

See what it does before how it is written

Elixir owns the API and validation. Erlang owns the queue and scheduler state. A full queue is rejected clearly.

Current code
Elixir
# The Elixir layer validates input and shapes public results
defmodule Scheduler.API do
  def submit(payload, opts \ []) do
    # A task is accepted only if validation and enqueueing both succeed
    with :ok <- validate(payload),
         {:ok, id} <- :scheduler_core.enqueue(payload, opts) do
      {:accepted, id}
    else
      # Reject clearly when the queue is full instead of taking more memory
      {:error, :queue_full} -> {:rejected, :busy}
      {:error, reason} -> {:rejected, reason}
    end
  end
end
DESIGN DECISION

Why this shape?

Why write a guarantee ledger before drawing a process tree?

Choice

Capacity, acknowledgements, durability, and idempotency define the promise. A process tree is one tool for keeping it.

Cost and boundary

With only an in-memory queue, promise bounded retries while the cluster is running—not at-least-once delivery across a full VM restart.

A FAMILIAR POINT OF VIEW

Coming from Java, Python, or JavaScript

Java

Familiar starting point
Executors, durable queues, and batch frameworks.
What BEAM changes
Supervision restores processes; acknowledgements, persistence, and idempotency determine job guarantees.
False friend
A process restart does not mean a job was persisted.

Python

Familiar starting point
Celery/RQ acknowledgements, retries, and workers.
What BEAM changes
BEAM expresses worker lifecycle in the application, while storage and protocol still decide delivery guarantees.
False friend
At-least-once means the business may observe duplicates.

JavaScript

Familiar starting point
BullMQ, cloud queues, and Promise workers.
What BEAM changes
OTP owns recovery boundaries; the queue protocol owns job state.
False friend
Retry belongs with idempotency, backoff, and a final-failure destination.
LAB
Hands on

Run four failure drills

About 15–25 minutes

Creating a problem on purpose is called fault injection. First write the expected result for a worker exit, a bad message, a timeout, and a disconnection.

  1. 01

    Make a running worker exit and check the retry count and capacity count

  2. 02

    Send a message outside the agreement and confirm that the service continues while leaving a log or Telemetry event

  3. 03

    Let a task pass its timeout and inspect isolation, cancellation rules, and a late result

  4. 04

    Disconnect a remote node and inspect the stale marker and state after reconnection

Copy into the terminal and press Enter
# Run only fault_injection tests and show their progress
mix test --only fault_injection --trace
What you should see
  • Processes that must stay alive remain under a supervision tree
  • A failed task retries only within its limit and never loops forever
  • Queue length and running count eventually return to a consistent state
  • Every rejection, recovery, or final stop leaves evidence in structured logs or metrics
Break it on purpose

Remove the retry limit and keep failing a task. Watch a retry storm make the problem larger.

What this shows

For these four named surprises, the system follows its written rules to recover, reject, or stop retrying.

What this does not show yet

Four drills cannot cover every failure or prove exactly-once delivery, cross-node consistency, or no repeated outside action.

Meet the words

Key ideas in this code

01

at-least-once

A task is tried at least once and may run more than once. Keeping that promise across a full VM restart requires durable storage or a log, acknowledgements, and idempotency records. An in-memory queue cannot claim it.

02

bounded queue

A waiting queue with fixed capacity. When it is full, the system must wait, reject, or degrade.

03

runbook

Operating instructions: which logs and metrics to inspect, how to judge a problem, and which safe actions to take.

Name the shape in the code

Design patterns used here

01

Bounded Buffer

Use a finite queue and reject new jobs explicitly when it is full.

02

Idempotent Consumer

Make duplicate execution from at-least-once delivery safe.

03

Retry with Exponential Backoff

Cap retries and increase their delay so failures do not create a retry storm.

Think it through

Before automatically retrying a failed task, what should you decide first?

Your turn

Ship the core, then choose the bilingual extension

Build the core runner in either Elixir or Erlang. Gather its supervision tree, capacity plan, failure record, release, and one-page runbook. If you completed interoperability, add an adapter or worker in the other language.

Hint 1Take the first step

Name at least one thing this version does not guarantee. That is honest engineering, not lost points.

Hint 2Make it a little smaller

Besides a successful submit, demonstrate queue_full so the audience sees the system admit that it is too busy.

Hint 3You are close

When writing the runbook, begin with what a user would notice and work backward to the signals an operator needs.

Hint 4Work backward from the finish

Choose one success signal and write the smallest test for it. If the computer cannot show the result, rewrite the signal as something you can truly observe.

Ready to move on when
  • Concurrency, queue size, timeout, and retry all have explicit limits

  • The project distinguishes in-memory retries while running from a cross-restart guarantee that needs persistence and acknowledgements

  • Another student can start the release and perform one safe check from the runbook

  • Optional extension: both languages have real responsibilities and tests cross the adapter in both directions

Take these with you

Remember three things

  1. 1

    A reliable project states its invariants, capacity, and failure choices before arranging process structure.

  2. 2

    Retries can multiply outside effects, so idempotency, backoff, and an attempt limit belong together.

  3. 3

    Maintainability is not a few extra log lines at the end. Leave useful evidence in the message agreement and record it in the runbook.

Read a little more

Visit the original sources

Mix and OTPErlang ApplicationsElixir Releases
Station completeYou ran the experiment and thought through the answer. Save this station.