Station R1 · Optional reviewOpen the adventure map
00Set up your toolsSetup01The starting lineSetup02Processes and mailboxesR1An Elixir data pipelineOptional reviewR2Read ErlangOptional review03Messages and timeouts04OTP rules for messages05A supervision tree that restarts06Put a limit on concurrency07When a BEAM node disappearsX1Agree on shared BEAM termsBilingual extensionX2Two-language partnersBilingual extension08A reliable job runner
Home/BEAM mainline/Station R1
R1
LanguageBeginner explorationOptional reviewElixir

An Elixir data pipeline

Trim, filter, then group. Let each small function do one job well.

5checkpoints
About 6 hours · Try 5 sessionssplit it into sessions
QUESTION · ONE PROBLEM FOR THIS STATION

[Optional language review] How can we split a log cleaner that meets blank lines, unknown levels, and malformed input into independently testable stages?

Start at the scene

One WARN line stopped the whole pipeline

The input mixes blank lines, spaces, and WARN entries. The parser only knows INFO and ERROR, so it raises FunctionClauseError before the summary can run.

Evidence you can observe
  • The smallest input that reproduces the failure, including one WARN line
  • The intermediate result after each trim, filter, and parse stage
  • The failing test and the function clause named by the exception
Why this station matters

The problem it solves

Long functions easily collect too many jobs. First use patterns to recognize data, then split each change into a small function. A pipeline passes one result to the next step.

After this station

You will be able to

  • Use a pattern to recognize data, then choose a path with a guard and a multi-clause function

  • Explain that Enum finishes a collection now, while Stream works step by step when results are requested

  • Write a small, clear set of ExUnit tests for pure functions that do not read files or send messages

Before you start
  • Know that BEAM has small processes and mailboxes, and have seen pattern matching once
  • Know that a list keeps items in order and a map finds values by keys
One job, two ways to write it

See what it does before how it is written

Both versions clean, parse, and count the same data.

Current code
Elixir
# Turn a multiline log into a count for each level
defmodule LogSummary do
  def summarize(lines) do
    # Clean whitespace first, then parse each line
    lines
    |> Stream.map(&String.trim/1)
    |> Stream.reject(&(&1 == ""))
    |> Enum.map(&parse_line/1)
    |> Enum.frequencies_by(& &1.level)
  end

  # Use string prefixes to take out the level and message
  defp parse_line("ERROR " <> message),
    do: %{level: :error, message: message}

  defp parse_line("INFO " <> message),
    do: %{level: :info, message: message}
end
LAB
Hands on

Swap two steps

About 15–25 minutes

Swap trim and reject. Guess where a line of spaces will go, then run the same data.

  1. 01

    Prepare the Elixir string list [" INFO boot ", " ", "ERROR timeout"]

  2. 02

    First trim both ends and then remove empty strings. Record the result

  3. 03

    Move reject before trim, then run exactly the same input again

Copy into the terminal and press Enter
# Run every test and show each test name and duration
mix test --trace
What you should see
  • With trim before reject, the line that contains only spaces is removed
  • With reject before trim, that line is not empty yet and may later enter the parser
Break it on purpose

Pass in WARN slow, then remove the fallback clause. See which clues remain in the FunctionClauseError.

What this shows

The order of data-processing steps changes what a later pattern can recognize.

What this does not show yet

This does not show that Stream is faster than Enum. Data size and the way results are consumed both matter.

Meet the words

Key ideas in this code

01

Pattern matching

The left side is a shape to check against the data on the right. A match takes values out. A mismatch tries another path. Here, = does more than assignment.

02

Multi-clause function

One function can have several entrances. The program checks from top to bottom and uses the first clause whose pattern and guard both fit. Arity is the number of arguments.

03

Pipeline

|> passes the result on its left to the function on its right as the first argument. It makes steps easier to see, but does not split functions for you.

Name the shape in the code

Design patterns used here

01

Pipes and Filters

Give each stage one transformation, then pass its output to the next stage.

02

Single Responsibility

Make cleaning, parsing, and summarizing separate testable functions.

Think it through

When is Stream worth considering first?

Your turn

Clean a log

Build a Mix tool that reads a log, skips blank lines, counts three levels, and finds the three most common messages.

Hint 1Take the first step

Begin with a pure function that understands one log line. Keep file reading at the outside edge.

Hint 2Make it a little smaller

Do not silently throw away an unknown line. Keep the clue by returning {:error, line}.

Hint 3You are close

Prepare tests for spaced text, an unknown level, and an empty file.

Hint 4Work backward from the finish

Choose one success signal and write the smallest test for it. If the computer cannot show the result, rewrite the signal as something you can truly observe.

Ready to move on when
  • Log cleaning and file reading are separate steps

  • An unknown line leaves a clear error result

  • ExUnit tests cover normal, edge, and error cases

Take these with you

Remember three things

  1. 1

    Look at the shape of the data before deciding which steps it should pass through.

  2. 2

    A pipeline passes results forward. Small, clear functions are what truly make it readable.

  3. 3

    Stream delays work, but it is not a “faster Enum.” Choose by the data and the goal.

Read a little more

Visit the original sources

Elixir basic typesEnumExUnit
Station completeYou ran the experiment and thought through the answer. Save this station.