feat(stdlib): add Unicode.* — UTF-8 code points, graphemes, case mapping (#13)
All checks were successful
docs / build-and-deploy (push) Successful in 2s
All checks were successful
docs / build-and-deploy (push) Successful in 2s
Make Ludic text correct-by-default over UTF-8, so player names, translated UI, and chat behave for every language instead of counting bytes and splitting characters in half. The byte-oriented Text.* stays for speed; Unicode.* is the layer that understands code points and (approximately) grapheme clusters. - len / byte_len code points vs bytes — the two lengths, kept distinct - is_valid_utf8 strict validation of untrusted input - char_at / chars code-point access by index; chars() -> []int - upper / lower case mapping (ASCII + Latin-1) - truncate first n code points, never a half-character - grapheme_len user-perceived characters (approx UAX#29) Pure integer/byte IR over NUL-terminated buffers; C-free, no data-table blob. Decoding and validation cover the full UTF-8 range (overlong/surrogate/>10FFFF rejected). grapheme_len collapses combining marks, variation selectors, ZWJ sequences (family emoji), and regional-indicator flag pairs. Documented v1 scope: wider-script/locale case rules (Latin-Extended, Greek, Cyrillic, Turkish i, German ß) and NFC normalization are follow-ups. - examples/library/unicode.ludic: asserts the invariants across ASCII, Latin-1 (é round-trips through upper/lower), a decomposed "café" (5 code points, 4 graphemes), a ZWJ family emoji (5 code points, 1 grapheme), and a flag (2 regional indicators, 1 grapheme). Wired into `x test` (now 55 passed). - docs: a new Unicode section + 9 per-symbol pages clarifying byte vs code point vs grapheme; inventory updated; every fence passes check-docs; site builds. - seed regenerated; `x bootstrap-cfree` fixpoint holds. Closes #13 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
b205dd8dfd
commit
bad6cd1ac9
19 changed files with 11934 additions and 10468 deletions
13
docs/language/unicode/_section.md
Normal file
13
docs/language/unicode/_section.md
Normal file
|
|
@ -0,0 +1,13 @@
|
|||
---
|
||||
id: unicode
|
||||
title: Unicode
|
||||
order: 6
|
||||
---
|
||||
|
||||
Correct-by-default text over UTF-8. Player names, translated menus, and chat arrive as UTF-8 bytes, and counting *bytes* gets non-ASCII text wrong — the wrong length, and truncation that slices a character in half into mojibake. The <a href="ns-Unicode"><code>Unicode</code></a> namespace works in <strong>code points</strong> and (approximately) <strong>grapheme clusters</strong> instead, so measuring, indexing, truncating, and case-mapping behave for every language.
|
||||
|
||||
This is the correctness layer, not a replacement: the byte-oriented <a href="ns-Text"><code>Text</code></a> operations stay for speed on ASCII and for raw byte work. Reach for <code>Unicode.*</code> whenever the text came from a human — a name, a message, a localized string.
|
||||
|
||||
Three notions of "length" matter, and the API keeps them distinct: <a href="unicode-byte_len"><code>byte_len</code></a> (storage), <a href="unicode-len"><code>len</code></a> (code points — Unicode scalar values), and <a href="unicode-grapheme_len"><code>grapheme_len</code></a> (user-perceived characters, where a base letter plus its combining accent, or a ZWJ emoji sequence, count as one).
|
||||
|
||||
**Coverage (v1):** decoding and <a href="unicode-is_valid_utf8"><code>validation</code></a> cover the full UTF-8 range. Case mapping covers ASCII and the Latin-1 letters — correct for Western-European text; wider scripts (Latin-Extended, Greek, Cyrillic), locale rules (Turkish <code>i</code>, German <code>ß</code>), and NFC normalization are follow-ups. <a href="unicode-grapheme_len"><code>grapheme_len</code></a> approximates UAX#29 for the cases real player text hits — combining marks, variation selectors, ZWJ sequences, and flag pairs.
|
||||
Loading…
Add table
Add a link
Reference in a new issue