feat(stdlib): add Unicode.* — UTF-8 code points, graphemes, case mapping (#13)
All checks were successful
docs / build-and-deploy (push) Successful in 2s
All checks were successful
docs / build-and-deploy (push) Successful in 2s
Make Ludic text correct-by-default over UTF-8, so player names, translated UI, and chat behave for every language instead of counting bytes and splitting characters in half. The byte-oriented Text.* stays for speed; Unicode.* is the layer that understands code points and (approximately) grapheme clusters. - len / byte_len code points vs bytes — the two lengths, kept distinct - is_valid_utf8 strict validation of untrusted input - char_at / chars code-point access by index; chars() -> []int - upper / lower case mapping (ASCII + Latin-1) - truncate first n code points, never a half-character - grapheme_len user-perceived characters (approx UAX#29) Pure integer/byte IR over NUL-terminated buffers; C-free, no data-table blob. Decoding and validation cover the full UTF-8 range (overlong/surrogate/>10FFFF rejected). grapheme_len collapses combining marks, variation selectors, ZWJ sequences (family emoji), and regional-indicator flag pairs. Documented v1 scope: wider-script/locale case rules (Latin-Extended, Greek, Cyrillic, Turkish i, German ß) and NFC normalization are follow-ups. - examples/library/unicode.ludic: asserts the invariants across ASCII, Latin-1 (é round-trips through upper/lower), a decomposed "café" (5 code points, 4 graphemes), a ZWJ family emoji (5 code points, 1 grapheme), and a flag (2 regional indicators, 1 grapheme). Wired into `x test` (now 55 passed). - docs: a new Unicode section + 9 per-symbol pages clarifying byte vs code point vs grapheme; inventory updated; every fence passes check-docs; site builds. - seed regenerated; `x bootstrap-cfree` fixpoint holds. Closes #13 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
b205dd8dfd
commit
bad6cd1ac9
19 changed files with 11934 additions and 10468 deletions
|
|
@ -457,5 +457,16 @@
|
|||
"os-config_dir",
|
||||
"os-cache_dir",
|
||||
"os-temp_dir"
|
||||
],
|
||||
"unicode": [
|
||||
"unicode-len",
|
||||
"unicode-byte_len",
|
||||
"unicode-is_valid_utf8",
|
||||
"unicode-char_at",
|
||||
"unicode-chars",
|
||||
"unicode-upper",
|
||||
"unicode-lower",
|
||||
"unicode-truncate",
|
||||
"unicode-grapheme_len"
|
||||
]
|
||||
}
|
||||
|
|
|
|||
|
|
@ -33,6 +33,7 @@ function selfhost_frags() -> []pointer {
|
|||
push(f, "selfhost/emit_noise.ludic")
|
||||
push(f, "selfhost/emit_log.ludic")
|
||||
push(f, "selfhost/emit_os.ludic")
|
||||
push(f, "selfhost/emit_unicode.ludic")
|
||||
push(f, "selfhost/emit_list.ludic")
|
||||
push(f, "selfhost/emit_ease.ludic")
|
||||
push(f, "selfhost/emit_collide.ludic")
|
||||
|
|
|
|||
|
|
@ -104,6 +104,7 @@ function cmd_test() -> int {
|
|||
feat_case("library/noise", "", "1 2 3 4 5 6 7 8 9 10 11", "noise.ludic (Noise value/perlin/simplex/fbm/cellular determinism + range)")
|
||||
feat_case("library/logging", "", "0 5 2 1", "logging.ludic (Log levels, set_level/level threshold, structured fields)")
|
||||
feat_case("library/os", "", "1 2 3 4 5 6 7 8 9 10 11 12 13", "os.ludic (Os args/env round-trip, platform/arch, known dirs)")
|
||||
feat_case("library/unicode", "", "1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18", "unicode.ludic (UTF-8 len/validate/char_at/chars/case/truncate/graphemes)")
|
||||
|
||||
# issue #9: the Time/Date/Duration/Clock stdlib, driven from its own `entry`.
|
||||
net_case("lang/offline_rewards", "13 650 2026-08-30 0")
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue