Unicode and UTF-8, Explained

Text feels like the easy part of programming right up until it is not. A user signs up with the name José, or pastes an address in Japanese, or drops an emoji into a form, and suddenly your clean strings turn into question marks and garbage like José. These bugs come from a confusion between two separate questions: which characters exist, and how do you store them as bytes. Unicode answered the first question. UTF-8 answered the second. Keeping them straight is most of the battle.

The mess before Unicode

In the beginning there was ASCII, which assigned numbers to 128 characters using 7 bits[1]. That covered the English alphabet, digits, and basic punctuation, and it was enough for American computing in the 1960s. It was not enough for the rest of the world.

What followed was a sprawl of incompatible encodings. Latin-1 and Windows-1252 covered Western European accents[2]. Different code pages covered Cyrillic, Greek, and Hebrew. Shift-JIS and others covered Japanese. The fatal problem was that they reused the same byte values for different characters. The byte 0xE9 might be é in one encoding and something else entirely in another. A document only made sense if you already knew which encoding it was written in, and that information was rarely attached to the file. Open it under the wrong assumption and you got nonsense.

Unicode: one number per character

Unicode is a single catalog that assigns a unique number, called a code point, to every character[3] in every writing system. The letter A is U+0041. The letter é is U+00E9. A grinning face emoji is U+1F600. The "U+" prefix just means it is a Unicode code point written in hexadecimal.

Unicode has room for over a million code points and has assigned around 150,000 of them so far[4], spanning living languages, historical scripts, symbols, and emoji. The key thing to understand is what Unicode is not. It is a numbering scheme, not a storage format. It says that é is U+00E9. It does not say how to write that as bytes on disk. That is a separate decision, and it is where encodings come in.

Encodings turn code points into bytes

An encoding is the rule for converting code points into the actual bytes you save or send. Several exist.

UTF-32 uses a fixed four bytes for every character[5]. It is simple to reason about because every character is the same size, but it wastes space, since most text does not need four bytes per character. UTF-16 uses two bytes for common characters and four for the rest. UTF-8 uses one to four bytes depending on the character[6], and it is the one that won.

Why UTF-8 won

UTF-8 has a property that made adoption almost painless: it is backward compatible with ASCII[7]. The first 128 code points encode as a single byte with exactly the same value they had in ASCII. That means every ASCII file ever written is already valid UTF-8, and plain English text takes one byte per character just as before. You pay for the wider character set only when you use it.

It is also variable width in a self-describing way. The leading bits of each byte say whether it starts a one, two, three, or four byte sequence, so a program can always find where the next character begins. It has no byte-order ambiguity, unlike UTF-16[8], which needed a special marker to say which end came first. These properties are why UTF-8 now covers around 98 percent of all web pages. When in doubt, the correct answer is almost always UTF-8.

Surrogate pairs and why emoji break length

UTF-16 could not fit code points above U+FFFF into its normal two bytes, so it uses a trick called surrogate pairs, representing one high code point as two sixteen-bit units[9]. This leaks into languages whose strings are built on UTF-16, JavaScript among them.

"😀".length        // 2   one character, but two UTF-16 units
[..."😀"].length   // 1   iterating yields one code point

This is the source of a whole family of bugs where an emoji counts as two, a slice cuts a character in half, or a reversed string comes out corrupted. The fix is to operate on code points rather than code units when you handle real user text.

Mojibake and the classic bugs

The garbled text you have seen, José instead of José, has a name: mojibake. It happens when bytes written in one encoding are read under another. The é was stored correctly as two UTF-8 bytes, but something along the way interpreted those two bytes as two separate Latin-1 characters and displayed them as é. The data was never lost. It was just read with the wrong rule.

Most of these problems vanish when you declare and use UTF-8 everywhere, end to end. Set your database and columns to UTF-8. Send charset=utf-8 in your HTTP headers[10]. Put <meta charset="utf-8"> at the top of your HTML[11]. Read and write files as UTF-8 explicitly rather than trusting the system default. The bugs come from mismatches, so removing the mismatch removes the bugs.

The clean mental model is two layers. Unicode is the agreement about which characters exist and what number each one gets. UTF-8 is the clever encoding that turns those numbers into bytes while staying compatible with the ASCII world that came before it. Hold those two ideas apart and most text problems become easy to diagnose. Collapse them into one vague notion of "text" and you will keep fighting the same ghost.

Sources (11)
  1. Wikipedia: ASCII
  2. Wikipedia: Windows-1252
  3. The Unicode Consortium: What is Unicode?
  4. Wikipedia: Unicode
  5. Wikipedia: UTF-32
  6. RFC 3629: UTF-8, a transformation format of ISO 10646
  7. Wikipedia: UTF-8
  8. Wikipedia: Comparison of Unicode encodings
  9. Wikipedia: UTF-16
  10. W3C: Declaring character encodings in HTTP
  11. WHATWG HTML Standard: The meta element