Elixir basics · Lesson 12 (open contents)
Count visible text and bytes
Compare graphemes, code points, bytes, and a simple Unicode-aware sigil.
Run this first
Run it once, change one input, and compare the new result.
# Compare visible text with its UTF-8 bytes
text = "长安🙂"
{String.length(text), byte_size(text), String.codepoints(text), Regex.match?(~r/安/u, text)}- The result is
{3, 10, ["长", "安", "🙂"], true}.
Read from the first line down
- Read
String.length/1: Count graphemes in a string. - Read
byte_size/1: Count bytes in a binary. - Read
~r/.../u: Build a Unicode-aware regular expression.
Symbols are not secret signs
String.length/1Count graphemes in a string.
byte_size/1Count bytes in a binary.
~r/.../uBuild a Unicode-aware regular expression.
Match each name to its meaning
grapheme
A grapheme is one user-visible character, even when several code points build it.
code point
A code point is one Unicode number; it is not always one visible character.
sigil
A sigil begins with ~ and gives a compact notation for text, regexes, dates, and other values.
Why these forms are useful
Run the example first. Predict one result, then change one input and run it again.
A grapheme is one user-visible character, even when several code points build it.
A code point is one Unicode number; it is not always one visible character.
Close the answer and try
Split Go🙂 into graphemes and count its bytes.
text = "Go🙂"
{String.____(text), ____(text)}Target result: Return {["G", "o", "🙂"], 6}.
Stuck? Read one hint
Use graphemes and byte_size.
After you run it, see one answer
text = "Go🙂"
{String.graphemes(text), byte_size(text)}Think about it: Must grapheme count equal byte count?
No. UTF-8 characters can use several bytes.
Remember these three lines
- 1
Use graphemes for visible text.
- 2
Use bytes for binary sizes and protocols.
- 3
Sigils are compact constructors, not a new data type.