--- id: unicode title: Unicode order: 6 --- Correct-by-default text over UTF-8. Player names, translated menus, and chat arrive as UTF-8 bytes, and counting *bytes* gets non-ASCII text wrong — the wrong length, and truncation that slices a character in half into mojibake. The Unicode namespace works in code points and (approximately) grapheme clusters instead, so measuring, indexing, truncating, and case-mapping behave for every language. This is the correctness layer, not a replacement: the byte-oriented Text operations stay for speed on ASCII and for raw byte work. Reach for Unicode.* whenever the text came from a human — a name, a message, a localized string. Three notions of "length" matter, and the API keeps them distinct: byte_len (storage), len (code points — Unicode scalar values), and grapheme_len (user-perceived characters, where a base letter plus its combining accent, or a ZWJ emoji sequence, count as one). **Coverage (v1):** decoding and validation cover the full UTF-8 range. Case mapping covers ASCII and the Latin-1 letters — correct for Western-European text; wider scripts (Latin-Extended, Greek, Cyrillic), locale rules (Turkish i, German ß), and NFC normalization are follow-ups. grapheme_len approximates UAX#29 for the cases real player text hits — combining marks, variation selectors, ZWJ sequences, and flag pairs.