ludic/docs/language/unicode/_section.md
Orkuncakilkaya bad6cd1ac9
All checks were successful
docs / build-and-deploy (push) Successful in 2s
feat(stdlib): add Unicode.* — UTF-8 code points, graphemes, case mapping (#13)
Make Ludic text correct-by-default over UTF-8, so player names, translated UI,
and chat behave for every language instead of counting bytes and splitting
characters in half. The byte-oriented Text.* stays for speed; Unicode.* is the
layer that understands code points and (approximately) grapheme clusters.

  - len / byte_len          code points vs bytes — the two lengths, kept distinct
  - is_valid_utf8           strict validation of untrusted input
  - char_at / chars         code-point access by index; chars() -> []int
  - upper / lower           case mapping (ASCII + Latin-1)
  - truncate                first n code points, never a half-character
  - grapheme_len            user-perceived characters (approx UAX#29)

Pure integer/byte IR over NUL-terminated buffers; C-free, no data-table blob.
Decoding and validation cover the full UTF-8 range (overlong/surrogate/>10FFFF
rejected). grapheme_len collapses combining marks, variation selectors, ZWJ
sequences (family emoji), and regional-indicator flag pairs. Documented v1
scope: wider-script/locale case rules (Latin-Extended, Greek, Cyrillic, Turkish
i, German ß) and NFC normalization are follow-ups.

- examples/library/unicode.ludic: asserts the invariants across ASCII, Latin-1
  (é round-trips through upper/lower), a decomposed "café" (5 code points, 4
  graphemes), a ZWJ family emoji (5 code points, 1 grapheme), and a flag (2
  regional indicators, 1 grapheme). Wired into `x test` (now 55 passed).
- docs: a new Unicode section + 9 per-symbol pages clarifying byte vs code point
  vs grapheme; inventory updated; every fence passes check-docs; site builds.
- seed regenerated; `x bootstrap-cfree` fixpoint holds.

Closes #13

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-30 22:29:00 +03:00

1.7 KiB

id title order
unicode Unicode 6

Correct-by-default text over UTF-8. Player names, translated menus, and chat arrive as UTF-8 bytes, and counting bytes gets non-ASCII text wrong — the wrong length, and truncation that slices a character in half into mojibake. The Unicode namespace works in code points and (approximately) grapheme clusters instead, so measuring, indexing, truncating, and case-mapping behave for every language.

This is the correctness layer, not a replacement: the byte-oriented Text operations stay for speed on ASCII and for raw byte work. Reach for Unicode.* whenever the text came from a human — a name, a message, a localized string.

Three notions of "length" matter, and the API keeps them distinct: byte_len (storage), len (code points — Unicode scalar values), and grapheme_len (user-perceived characters, where a base letter plus its combining accent, or a ZWJ emoji sequence, count as one).

Coverage (v1): decoding and validation cover the full UTF-8 range. Case mapping covers ASCII and the Latin-1 letters — correct for Western-European text; wider scripts (Latin-Extended, Greek, Cyrillic), locale rules (Turkish i, German ß), and NFC normalization are follow-ups. grapheme_len approximates UAX#29 for the cases real player text hits — combining marks, variation selectors, ZWJ sequences, and flag pairs.