Erlang basics · Lesson 12 (open contents)
01 · Open erl02 · Meet common terms03 · Tell two kinds of text apart04 · Bind a variable once05 · Split and filter a list06 · Parse a number safely07 · Choose one path with case08 · Let a function keep going09 · Put code in a module10 · Compare terms and Boolean results11 · Read and update nested maps12 · Build and inspect UTF-8 binaries13 · Let clauses choose the function path14 · Solve a list twice15 · Choose with patterns, guards, and errors16 · Describe a module and one record17 · Put one tested module in Rebar318 · Return success or failure as data19 · Read and write one file safely20 · Build a data pipeline with functions21 · Keep records behind a module boundary22 · Define a module contract23 · Describe the data contract24 · Test small rules and whole flows25 · Shape builds and make a command26 · Write the project promise27 · Parse one real line28 · Keep the public API small29 · Make the promise executable30 · Package the edge and leave clues31 · Prove the project is done
ERLANG · FOUNDATION · LESSON 1225 minutes

Build and inspect UTF-8 binaries

Compare code points, bytes, Unicode conversion, and bit-syntax segments.

01 · Start with the whole example

Run this first

Run it once, change one input, and compare the new result.

Erlang
%% Inspect Unicode as code points and bytes
Text = unicode:characters_to_binary([16#957F, 16#5B89, 16#1F642]),
<<First/utf8, Rest/binary>> = Text,
{First, byte_size(Text), Rest}.
Check the result first
  • First is code point 16#957F; the full binary uses 10 bytes.
02 · Take the code apart

Read from the first line down

  1. Read <<Head/utf8, Rest/binary>>: Read one UTF-8 code point and keep the remaining bytes.
  2. Read byte_size/1: Count bytes in a binary.
  3. Read unicode:characters_to_binary/1: Convert character data to UTF-8 bytes.
03 · Meet the new symbols

Symbols are not secret signs

<<Head/utf8, Rest/binary>>

Read one UTF-8 code point and keep the remaining bytes.

byte_size/1

Count bytes in a binary.

unicode:characters_to_binary/1

Convert character data to UTF-8 bytes.

04 · Ideas inside the code

Match each name to its meaning

01

bit syntax

<<>> builds binaries from integer, binary, and typed segments.

02

UTF-8 segment

Codepoint/utf8 encodes one Unicode code point into its UTF-8 bytes.

03

Unicode conversion

unicode:characters_to_binary/1 converts supported character data into a UTF-8 binary.

05 · Make the idea clear

Why these forms are useful

Run the example first. Predict one result, then change one input and run it again.

<<>> builds binaries from integer, binary, and typed segments.

Codepoint/utf8 encodes one Unicode code point into its UTF-8 bytes.

06 · Change it yourself

Close the answer and try

Build the UTF-8 binary for A🙂 and read its first code point.

Practice starting point
Text = unicode:characters_to_binary([$A, 16#1F642]),
<<First/____, Rest/____>> = Text,
{First, Rest}.

Target result: Use utf8 and binary; First should be 65.

Stuck? Read one hint

The first segment is one code point, the rest is bytes.

After you run it, see one answer
One answer
Text = unicode:characters_to_binary([$A, 16#1F642]),
<<First/utf8, Rest/binary>> = Text,
{First, Rest}.
Think about it: Is byte count the same as visible-character count?

No. A UTF-8 code point can occupy several bytes, and graphemes can use several code points.

Take with you

Remember these three lines

  1. 1

    Use binaries for modern text boundaries.

  2. 2

    State UTF-8 in bit syntax.

  3. 3

    Measure bytes only when bytes are the real unit.

Lesson completeRun the code once, then mark this lesson as done.