Split the flat 38-file selfhost/ into concern-based subdirectories:
frontend/ lex, parse, parse_game, ast
support/ str, buf, io
backend/ core IR + expression/statement lowering
backend/game/ ECS/scene/event/world lowering
backend/stdlib/ the namespaced Math.*/Text.*/Crypto.*/… intrinsics
and split the three oversized emitters at responsibility boundaries so
no file mixes concerns:
emit_game.ludic -> + emit_world.ludic (reflection world table,
tick helpers, @main synthesis)
emit_expr.ludic -> + emit_call.ludic (namespaced builtins, call
lowering, expr dispatch)
emit_text.ludic -> + emit_text_prelude.ludic (emitted string-builder runtime)
FRAGS in tools/x/selfhost.ludic is updated to the new paths with the link
order preserved, and the Python doc/vocabulary tooling is updated to walk
the new layout. Because the build is a plain in-order concatenation and
every split lands on a blank-line boundary, the regenerated seed is
byte-identical: `x reseed` leaves selfhost/ludicc.seed.ll unchanged,
`x bootstrap-cfree` still reaches its fixed point, and both `x test` (56)
and `x selfhost-test` (29, incl. golden renders) stay green.
Closes#29
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Make Ludic text correct-by-default over UTF-8, so player names, translated UI,
and chat behave for every language instead of counting bytes and splitting
characters in half. The byte-oriented Text.* stays for speed; Unicode.* is the
layer that understands code points and (approximately) grapheme clusters.
- len / byte_len code points vs bytes — the two lengths, kept distinct
- is_valid_utf8 strict validation of untrusted input
- char_at / chars code-point access by index; chars() -> []int
- upper / lower case mapping (ASCII + Latin-1)
- truncate first n code points, never a half-character
- grapheme_len user-perceived characters (approx UAX#29)
Pure integer/byte IR over NUL-terminated buffers; C-free, no data-table blob.
Decoding and validation cover the full UTF-8 range (overlong/surrogate/>10FFFF
rejected). grapheme_len collapses combining marks, variation selectors, ZWJ
sequences (family emoji), and regional-indicator flag pairs. Documented v1
scope: wider-script/locale case rules (Latin-Extended, Greek, Cyrillic, Turkish
i, German ß) and NFC normalization are follow-ups.
- examples/library/unicode.ludic: asserts the invariants across ASCII, Latin-1
(é round-trips through upper/lower), a decomposed "café" (5 code points, 4
graphemes), a ZWJ family emoji (5 code points, 1 grapheme), and a flag (2
regional indicators, 1 grapheme). Wired into `x test` (now 55 passed).
- docs: a new Unicode section + 9 per-symbol pages clarifying byte vs code point
vs grapheme; inventory updated; every fence passes check-docs; site builds.
- seed regenerated; `x bootstrap-cfree` fixpoint holds.
Closes#13
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>