Proposal: Unicode-aware text (Unicode.*) — code points, graphemes, case mapping for localization #13

Closed
opened 2026-08-29 20:21:57 +02:00 by orkun · 0 comments
Owner

Summary

Make Ludic text truly Unicode-aware: correct handling of UTF-8, code
points, grapheme clusters, and case mapping — so player names, translated UI,
and chat behave correctly for every language.

Why it matters for game devs

  • Localization: menus and dialogue in Turkish, Japanese, emoji, RTL scripts.
  • Player-entered text: names, chat, level titles — measured and truncated
    without cutting a character in half.
  • Text rendering: laying out glyphs needs correct code-point / grapheme
    iteration, not byte counting.

Current gap

Text.* exists but is largely byte-oriented (Text.length, Text.char_at,
Text.slice). For non-ASCII this returns wrong lengths and can split multi-byte
characters, producing mojibake.

Proposed API (illustrative)

# doc-check: skip — illustrative API sketch
Unicode.len("héllo")          # 5 code points (not 6 bytes)
Unicode.grapheme_len("👨‍👩‍👧") # 1 user-perceived character
for cp in Unicode.chars(name) { draw_glyph(cp) }
Unicode.upper("straße")       # locale-aware case mapping
Unicode.is_valid_utf8(bytes)  # validate untrusted input (chat, files)
  • Code-point and grapheme-cluster iteration and length.
  • Correct case folding (upper/lower/casefold).
  • UTF-8 validation + safe truncation to a display width.
  • Optional: normalization (NFC) for comparing names.

Considerations

  • Ship a compact Unicode data table (case + grapheme break); full ICU is too
    heavy. Document the Unicode version supported.
  • Native/C-free; table generated at build time.
  • Byte-level Text.* stays for performance; Unicode API is the correct-by-default
    layer. Clarify which is which in docs.
  • Ties into text rendering and the Filesystem (reading UTF-8 files).

Scope / acceptance

  • Unicode namespace: chars/graphemes/len/case/validate/truncate.
  • Bundled, versioned, compact data tables.
  • Docs page clarifying byte vs code-point vs grapheme.
  • Tests across scripts + emoji ZWJ sequences.

Related: #2 (Text namespace), Filesystem & IO, Regex.

## Summary Make Ludic **text truly Unicode-aware**: correct handling of UTF-8, code points, grapheme clusters, and case mapping — so player names, translated UI, and chat behave correctly for every language. ## Why it matters for game devs - **Localization**: menus and dialogue in Turkish, Japanese, emoji, RTL scripts. - **Player-entered text**: names, chat, level titles — measured and truncated without cutting a character in half. - **Text rendering**: laying out glyphs needs correct code-point / grapheme iteration, not byte counting. ## Current gap `Text.*` exists but is largely **byte-oriented** (`Text.length`, `Text.char_at`, `Text.slice`). For non-ASCII this returns wrong lengths and can split multi-byte characters, producing mojibake. ## Proposed API (illustrative) ```ludic # doc-check: skip — illustrative API sketch Unicode.len("héllo") # 5 code points (not 6 bytes) Unicode.grapheme_len("👨‍👩‍👧") # 1 user-perceived character for cp in Unicode.chars(name) { draw_glyph(cp) } Unicode.upper("straße") # locale-aware case mapping Unicode.is_valid_utf8(bytes) # validate untrusted input (chat, files) ``` - Code-point and **grapheme-cluster** iteration and length. - Correct case folding (`upper`/`lower`/`casefold`). - UTF-8 validation + safe truncation to a display width. - Optional: normalization (NFC) for comparing names. ## Considerations - Ship a **compact** Unicode data table (case + grapheme break); full ICU is too heavy. Document the Unicode version supported. - Native/C-free; table generated at build time. - Byte-level `Text.*` stays for performance; Unicode API is the correct-by-default layer. Clarify which is which in docs. - Ties into text rendering and the Filesystem (reading UTF-8 files). ## Scope / acceptance - [ ] `Unicode` namespace: chars/graphemes/len/case/validate/truncate. - [ ] Bundled, versioned, compact data tables. - [ ] Docs page clarifying byte vs code-point vs grapheme. - [ ] Tests across scripts + emoji ZWJ sequences. Related: #2 (Text namespace), Filesystem & IO, Regex.
orkun added the
proposal
priority:medium
area:stdlib
labels 2026-08-29 20:21:57 +02:00
orkun closed this issue 2026-08-30 21:29:02 +02:00
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference: workshopsoft/ludic#13
No description provided.