Computing and the Command Line

Text as Numbers

Text is stored as numbers too. ASCII's 128 characters, one byte each; Unicode's code points for every writing system, written U+00E5; and UTF-8, which stores ASCII as before and everything else in two to four bytes, following a bit pattern you can read. Seeing the bytes with xxd and Python's ord, chr, and encode; why a string's length in characters and in bytes differ; the garbled text that appears when bytes are read with the wrong encoding; and LF versus CRLF line endings.

  • 6 min
  • 9 steps
  • 2 questions
  • Lesson 64 of 80

In this lesson

  1. Characters are numbers
  2. ASCII
  3. Unicode
  4. UTF-8
  5. Characters versus bytes
  6. When the encoding is wrong
  7. Line endings
  8. Your turn
  9. So

Characters are numbers

A computer can’t store the letter A, only bits. So text is stored as numbers, one (or more) per character, by agreement about which number means which character: a character encoding. Getting that agreement right is why text usually just works, and getting it wrong is why it sometimes turns into Ã¥.

Top, the characters of Hej då followed by a newline, each over its bytes in hex: H 48, e 65, j 6a, space 20, d 64, å c3 a5, newline 0a; seven characters, eight bytes. The command printf 'Hej då\n' | xxd prints 00000000: 4865 6a20 64c3 a50a, Hej d.... Middle, UTF-8 uses one to four bytes per code point: U+0000 to U+007F, 0xxxxxxx, A is 41, ASCII, one byte; U+0080 to U+07FF, 110xxxxx 10xxxxxx, å is c3 a5, two bytes; U+0800 to U+FFFF, 1110xxxx 10xxxxxx 10xxxxxx, the euro sign is e2 82 ac, three bytes; U+10000 to U+10FFFF, 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx, the fox emoji is f0 9f a6 8a, four bytes. Bottom, line endings: Linux, line followed by \n, bytes 6c 69 6e 65 0a; Windows, line followed by \r\n, bytes 6c 69 6e 65 0d 0a.
Characters are code points; UTF-8 turns each into one to four bytes. Credit: StudyCorner diagram · CC BY 4.0 · Source

ASCII

ASCII, the American Standard Code for Information Interchange, is the original agreement still underneath everything: 128 codes, 0 to 127, which fit in seven bits 1. 95 are printable, the digits, upper- and lowercase English letters, and punctuation; 33 are control codes from the days of Teletypes, of which newline and tab are the ones you’ll meet 1. A is 65, a is 97, and the digit 0 is 48.

xxd shows the bytes in a file or a pipe, in hex, with the text alongside:

me@linuxbox:~$ printf 'A' | xxd
00000000: 41                                       A
me@linuxbox:~$ printf 'Hi!\n' | xxd
00000000: 4869 210a                                Hi!.

41 is 65 in hex; 0a is 10, the newline, shown as . in the text column because it isn’t printable. Python converts between characters and numbers with ord and chr 2:

me@linuxbox:~$ python3 -c 'print(ord("A"), chr(97))'
65 a

Quick check

printf 'A' | xxd shows 41. What is that?

Unicode

ASCII has no å, no ö, no €, no Japanese, no emoji. Unicode assigns every character in every writing system a number called a code point, from 0 up to 0x10FFFF, about 1.1 million possible values, written like U+00E5 for å 3. Its first 128 code points are exactly ASCII 1.

me@linuxbox:~$ python3 -c 'print(ord("€"), hex(ord("€")))'
8364 0x20ac

The euro sign is code point U+20AC. A code point is just a number, though; to store it, it has to be turned into bytes, and that’s what an encoding like UTF-8 does.

UTF-8

UTF-8 is one of the most commonly used encodings 3, and the one everything in this course’s Linux tools expects. Its rules 3 4:

  • A code point below 128 (ASCII) is one byte, the same byte as in ASCII. So plain ASCII text is already valid UTF-8 3.
  • Anything higher becomes two, three, or four bytes, each between 128 and 255 3.

The bytes follow a pattern you can read 4:

Code points Bytes
U+0000 to U+007F 0xxxxxxx
U+0080 to U+07FF 110xxxxx 10xxxxxx
U+0800 to U+FFFF 1110xxxx 10xxxxxx 10xxxxxx
U+10000 to U+10FFFF 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx

The x bits hold the code point; the first byte’s leading 1s say how many bytes follow, and continuation bytes start with 10. Here’s å, U+00E5, in bits:

me@linuxbox:~$ printf 'å' | xxd -b
00000000: 11000011 10100101                                      ..

110 then 00011, and 10 then 100101: put the x bits together, 00011100101, and that’s 0xE5. Three- and four-byte characters work the same way:

me@linuxbox:~$ printf '€' | xxd
00000000: e282 ac                                  ...
me@linuxbox:~$ printf '🦊' | xxd
00000000: f09f a68a                                ....

Because every byte of a multi-byte character is 128 or more, UTF-8 never confuses part of a character with an ASCII one, and if bytes are lost, a reader can find the start of the next character again 3.

Quick check

Why is the UTF-8 encoding of Hej då one byte longer than its number of characters?

Characters versus bytes

So the length of text depends on how you count:

me@linuxbox:~$ printf 'Hej då\n' | xxd
00000000: 4865 6a20 64c3 a50a                      Hej d...
me@linuxbox:~$ python3 -c 'print(len("Hej då"), len("Hej då".encode("utf-8")))'
6 7

Six characters, seven bytes, because å takes two. .encode("utf-8") turns text into bytes 3:

me@linuxbox:~$ python3 -c 'print("å".encode("utf-8"))'
b'\xc3\xa5'

When the encoding is wrong

Bytes don’t say which encoding wrote them. Read UTF-8 bytes as if they were an older one-byte encoding, Latin-1, and each byte of å becomes its own character:

me@linuxbox:~$ python3 -c 'print("Hej då".encode("utf-8").decode("latin-1"))'
Hej då

That Ã¥ is the classic sign of UTF-8 text displayed with the wrong encoding, common in old emails and CSV files from older programs. The fix is to read the file as UTF-8, not to change the text. file guesses a file’s encoding 5:

me@linuxbox:~$ file utf8.txt
utf8.txt: Unicode text, UTF-8 text

Line endings

One more difference hides at the end of every line. Linux ends lines with a single LF, line feed, byte 0a, written \n. Windows ends them with CR LF, carriage return then line feed, 0d 0a, written \r\n:

me@linuxbox:~$ printf 'line\r\n' | xxd
00000000: 6c69 6e65 0d0a                           line..
me@linuxbox:~$ file crlf.txt
crlf.txt: ASCII text, with CRLF line terminators

A script saved on Windows with CRLF endings can fail on Linux with a confusing error, since the shell sees the invisible \r as part of each line. This is what git’s core.autocrlf setting and its line-ending warnings are about (Git course, module 1).

Your turn

Exercises

  1. printf 'Hello\n' | xxd. Find each letter’s byte; which is the newline?
  2. python3 -c 'print([ord(c) for c in "Linux"])'. Check two of them against xxd.
  3. printf 'ö' | xxd -b. Using the UTF-8 table, decode the bits back to the code point, and check with python3 -c 'print(hex(ord("ö")))'.
  4. How many bytes is Smörgåsbord in UTF-8? Count first, then check with printf 'Smörgåsbord' | wc -c.
  5. Make a file with CRLF endings (printf 'a\r\nb\r\n' > win.txt), check it with file, then convert it with tr -d '\r' < win.txt > unix.txt and check again.
  6. python3 -c 'print("Smörgåsbord".encode("utf-8").decode("latin-1"))'. What does mis-decoded text look like?
Answers
  1. 48 65 6c 6c 6f 0a: H, e, l, l, o, and the newline 0a.
  2. [76, 105, 110, 117, 120]; L is 76 = 0x4c.
  3. 11000011 10110110: the x bits are 00011 and 110110, together 00011110110 = 0xF6, so U+00F6.
  4. 13 bytes: 11 characters, two of them (ö and å) two bytes each.
  5. ASCII text, with CRLF line terminators, then plain ASCII text.
  6. Smörgåsbord.

So

Text is numbers. ASCII gives 128 characters a byte each; Unicode gives every character a code point, up to U+10FFFF; UTF-8 stores ASCII as one byte and everything else in two to four, following a pattern of leading bits you can decode by hand. xxd shows the bytes, Python’s ord, chr, and .encode convert, and file guesses an encoding. Lengths differ in characters and bytes, Ã¥ means UTF-8 read the wrong way, and Linux ends lines with LF where Windows uses CR LF.

Lesson complete

Nice work.

1day streak
0/1today's goal
–correct

Up next · 6 min

Integers, Overflow, and Floating Point

Next lesson
Sources for this lesson
  1. 1
    ASCII. Wikipedia. verifiedAmerican Standard Code for Information Interchange: 128 code points, 0-127, storable in seven bits; 95 printable characters and 33 control codes originating with Teletypes; the letter i is 105; the first 128 Unicode code points are the same as ASCII.
  2. 2
    Built-in Functions (Python documentation). Python Software Foundation. verifiedbin(), hex(), oct() convert an integer to a prefixed string; int(text, base) parses one; ord() gives a character's Unicode code point and chr() the reverse (chr(97) is 'a', chr(8364) is the euro sign). sum(): since 3.12, summation of floats uses an algorithm with higher accuracy. id() is, in CPython, the address of the object in memory. open() buffers binary files in fixed-size chunks by default; print()'s output buffering is set by the file, and flush=True forces it out.
  3. 3
    Unicode HOWTO (Python documentation). Python Software Foundation. verifiedCode points are integers from 0 to 0x10FFFF, written U+265E; UTF-8 is one of the most commonly used encodings: code points below 128 are one byte, others two to four bytes each between 128 and 255; ASCII text is valid UTF-8; embedded zero bytes only for U+0000; a reader can resynchronize after lost bytes; str.encode and bytes.decode convert.
  4. 4
    F. Yergeau. RFC 3629: UTF-8, a transformation format of ISO 10646. IETF. 2003. verifiedUTF-8 byte patterns by code point range: 0000-007F 0xxxxxxx; 0080-07FF 110xxxxx 10xxxxxx; 0800-FFFF 1110xxxx 10xxxxxx 10xxxxxx; 10000-10FFFF 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx; only one valid encoding per character.
  5. 5
    file(1) manual page. man7.org (Linux man-pages). verifiedfile determines file type by testing contents, including text encodings such as ASCII and UTF-8 and line terminators.