Station 08Open the adventure map
A reliable job runner
Protect capacity and failure behavior in one language first. Add a second-language boundary as an extension.
How do we keep capacity and retries bounded through queue-full, worker exit, timeout, and node loss—and state honestly what is still missing across a full VM restart?
One persistently failing job caused a retry storm
A failing job has no retry limit and no idempotency key. It returns to the queue again and again, capacity counters drift, and an external action may happen more than once.
- Queued, running, and retry counts before and after each fault
- The idempotency key and a record proving the external effect happened once
- Reject, recover, and give-up metrics plus injected-fault tests
The problem it solves
The final project joins the earlier tools. We will fill the queue, fail a task, delay a message, and disconnect a node, then check state, retries, recovery, and rejection rules.
You will be able to
Write down rules that must always hold before arranging processes, queues, and supervision trees
Ship a one-language core first, then decide whether an Erlang/Elixir adapter adds real value
Create small failures on purpose and check how the system recovers, rejects work, or leaves clues
- Finish the preceding BEAM mainline and keep the queue and supervision exercises
- The core needs one language; the bilingual extension also needs the shared-terms and interoperability stations
See what it does before how it is written
Elixir owns the API and validation. Erlang owns the queue and scheduler state. A full queue is rejected clearly.
# The Elixir layer validates input and shapes public results
defmodule Scheduler.API do
def submit(payload, opts \ []) do
# A task is accepted only if validation and enqueueing both succeed
with :ok <- validate(payload),
{:ok, id} <- :scheduler_core.enqueue(payload, opts) do
{:accepted, id}
else
# Reject clearly when the queue is full instead of taking more memory
{:error, :queue_full} -> {:rejected, :busy}
{:error, reason} -> {:rejected, reason}
end
end
endWhy this shape?
Why write a guarantee ledger before drawing a process tree?
Capacity, acknowledgements, durability, and idempotency define the promise. A process tree is one tool for keeping it.
With only an in-memory queue, promise bounded retries while the cluster is running—not at-least-once delivery across a full VM restart.
Coming from Java, Python, or JavaScript
Java
- Familiar starting point
- Executors, durable queues, and batch frameworks.
- What BEAM changes
- Supervision restores processes; acknowledgements, persistence, and idempotency determine job guarantees.
- False friend
- A process restart does not mean a job was persisted.
Python
- Familiar starting point
- Celery/RQ acknowledgements, retries, and workers.
- What BEAM changes
- BEAM expresses worker lifecycle in the application, while storage and protocol still decide delivery guarantees.
- False friend
- At-least-once means the business may observe duplicates.
JavaScript
- Familiar starting point
- BullMQ, cloud queues, and Promise workers.
- What BEAM changes
- OTP owns recovery boundaries; the queue protocol owns job state.
- False friend
- Retry belongs with idempotency, backoff, and a final-failure destination.
Run four failure drills
Creating a problem on purpose is called fault injection. First write the expected result for a worker exit, a bad message, a timeout, and a disconnection.
- 01
Make a running worker exit and check the retry count and capacity count
- 02
Send a message outside the agreement and confirm that the service continues while leaving a log or Telemetry event
- 03
Let a task pass its timeout and inspect isolation, cancellation rules, and a late result
- 04
Disconnect a remote node and inspect the stale marker and state after reconnection
# Run only fault_injection tests and show their progress
mix test --only fault_injection --trace- Processes that must stay alive remain under a supervision tree
- A failed task retries only within its limit and never loops forever
- Queue length and running count eventually return to a consistent state
- Every rejection, recovery, or final stop leaves evidence in structured logs or metrics
Remove the retry limit and keep failing a task. Watch a retry storm make the problem larger.
For these four named surprises, the system follows its written rules to recover, reject, or stop retrying.
Four drills cannot cover every failure or prove exactly-once delivery, cross-node consistency, or no repeated outside action.
Key ideas in this code
at-least-once
A task is tried at least once and may run more than once. Keeping that promise across a full VM restart requires durable storage or a log, acknowledgements, and idempotency records. An in-memory queue cannot claim it.
bounded queue
A waiting queue with fixed capacity. When it is full, the system must wait, reject, or degrade.
runbook
Operating instructions: which logs and metrics to inspect, how to judge a problem, and which safe actions to take.
Design patterns used here
Bounded Buffer
Use a finite queue and reject new jobs explicitly when it is full.
Idempotent Consumer
Make duplicate execution from at-least-once delivery safe.
Retry with Exponential Backoff
Cap retries and increase their delay so failures do not create a retry storm.
Before automatically retrying a failed task, what should you decide first?
Ship the core, then choose the bilingual extension
Build the core runner in either Elixir or Erlang. Gather its supervision tree, capacity plan, failure record, release, and one-page runbook. If you completed interoperability, add an adapter or worker in the other language.
Hint 1Take the first step
Name at least one thing this version does not guarantee. That is honest engineering, not lost points.
Hint 2Make it a little smaller
Besides a successful submit, demonstrate queue_full so the audience sees the system admit that it is too busy.
Hint 3You are close
When writing the runbook, begin with what a user would notice and work backward to the signals an operator needs.
Hint 4Work backward from the finish
Choose one success signal and write the smallest test for it. If the computer cannot show the result, rewrite the signal as something you can truly observe.
Concurrency, queue size, timeout, and retry all have explicit limits
The project distinguishes in-memory retries while running from a cross-restart guarantee that needs persistence and acknowledgements
Another student can start the release and perform one safe check from the runbook
Optional extension: both languages have real responsibilities and tests cross the adapter in both directions
Remember three things
- 1
A reliable project states its invariants, capacity, and failure choices before arranging process structure.
- 2
Retries can multiply outside effects, so idempotency, backoff, and an attempt limit belong together.
- 3
Maintainability is not a few extra log lines at the end. Leave useful evidence in the message agreement and record it in the runbook.