feat(stdlib): add Unicode.* — UTF-8 code points, graphemes, case mapping (#13)
All checks were successful
docs / build-and-deploy (push) Successful in 2s

Make Ludic text correct-by-default over UTF-8, so player names, translated UI,
and chat behave for every language instead of counting bytes and splitting
characters in half. The byte-oriented Text.* stays for speed; Unicode.* is the
layer that understands code points and (approximately) grapheme clusters.

  - len / byte_len          code points vs bytes — the two lengths, kept distinct
  - is_valid_utf8           strict validation of untrusted input
  - char_at / chars         code-point access by index; chars() -> []int
  - upper / lower           case mapping (ASCII + Latin-1)
  - truncate                first n code points, never a half-character
  - grapheme_len            user-perceived characters (approx UAX#29)

Pure integer/byte IR over NUL-terminated buffers; C-free, no data-table blob.
Decoding and validation cover the full UTF-8 range (overlong/surrogate/>10FFFF
rejected). grapheme_len collapses combining marks, variation selectors, ZWJ
sequences (family emoji), and regional-indicator flag pairs. Documented v1
scope: wider-script/locale case rules (Latin-Extended, Greek, Cyrillic, Turkish
i, German ß) and NFC normalization are follow-ups.

- examples/library/unicode.ludic: asserts the invariants across ASCII, Latin-1
  (é round-trips through upper/lower), a decomposed "café" (5 code points, 4
  graphemes), a ZWJ family emoji (5 code points, 1 grapheme), and a flag (2
  regional indicators, 1 grapheme). Wired into `x test` (now 55 passed).
- docs: a new Unicode section + 9 per-symbol pages clarifying byte vs code point
  vs grapheme; inventory updated; every fence passes check-docs; site builds.
- seed regenerated; `x bootstrap-cfree` fixpoint holds.

Closes #13

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Orkun ÇAKILKAYA 2026-08-30 22:29:00 +03:00
parent b205dd8dfd
commit bad6cd1ac9
19 changed files with 11934 additions and 10468 deletions

View file

@ -0,0 +1,13 @@
---
id: unicode
title: Unicode
order: 6
---
Correct-by-default text over UTF-8. Player names, translated menus, and chat arrive as UTF-8 bytes, and counting *bytes* gets non-ASCII text wrong — the wrong length, and truncation that slices a character in half into mojibake. The <a href="ns-Unicode"><code>Unicode</code></a> namespace works in <strong>code points</strong> and (approximately) <strong>grapheme clusters</strong> instead, so measuring, indexing, truncating, and case-mapping behave for every language.
This is the correctness layer, not a replacement: the byte-oriented <a href="ns-Text"><code>Text</code></a> operations stay for speed on ASCII and for raw byte work. Reach for <code>Unicode.*</code> whenever the text came from a human — a name, a message, a localized string.
Three notions of "length" matter, and the API keeps them distinct: <a href="unicode-byte_len"><code>byte_len</code></a> (storage), <a href="unicode-len"><code>len</code></a> (code points — Unicode scalar values), and <a href="unicode-grapheme_len"><code>grapheme_len</code></a> (user-perceived characters, where a base letter plus its combining accent, or a ZWJ emoji sequence, count as one).
**Coverage (v1):** decoding and <a href="unicode-is_valid_utf8"><code>validation</code></a> cover the full UTF-8 range. Case mapping covers ASCII and the Latin-1 letters — correct for Western-European text; wider scripts (Latin-Extended, Greek, Cyrillic), locale rules (Turkish <code>i</code>, German <code>ß</code>), and NFC normalization are follow-ups. <a href="unicode-grapheme_len"><code>grapheme_len</code></a> approximates UAX#29 for the cases real player text hits — combining marks, variation selectors, ZWJ sequences, and flag pairs.

View file

@ -0,0 +1,25 @@
---
id: unicode-byte_len
name: Unicode.byte_len
category: unicode
kind: namespace-method
tokens: Unicode.byte_len
sig: Unicode.byte_len(s) -> int
tip: Number of bytes in a string.
order: 2
ns: Unicode
member: byte_len
---
The number of <strong>bytes</strong> the string occupies — its storage size. This is what raw byte APIs and buffers care about, and it is greater than <a href="unicode-len"><code>len</code></a> whenever the text contains non-ASCII code points (each of which encodes to 2–4 bytes).
Parameters:
- `s` — the UTF-8 string
```ludic
program ByteLen {
entry {
print(Unicode.byte_len("hello")) # 5 (all ASCII)
}
}
```

View file

@ -0,0 +1,26 @@
---
id: unicode-char_at
name: Unicode.char_at
category: unicode
kind: namespace-method
tokens: Unicode.char_at
sig: Unicode.char_at(s, i) -> int
tip: The i-th code point of a string.
order: 4
ns: Unicode
member: char_at
---
Returns the code point at code-point index <code>i</code> as an integer (a Unicode scalar value), or <code>-1</code> when <code>i</code> is out of range. Indexing is by code point, not byte, so it never lands in the middle of a multibyte character.
Parameters:
- `s` — the UTF-8 string
- `i` — the code-point index (0-based)
```ludic
program CharAt {
entry {
print(Unicode.char_at("hello", 0)) # 104 = 'h'
}
}
```

View file

@ -0,0 +1,27 @@
---
id: unicode-chars
name: Unicode.chars
category: unicode
kind: namespace-method
tokens: Unicode.chars
sig: Unicode.chars(s) -> []int
tip: Every code point of a string, in order.
order: 5
ns: Unicode
member: chars
---
Decodes the whole string into a <a href="type-slice"><code>[]int</code></a> of code points, in order — the iteration primitive for laying out glyphs, scanning text, or transforming character by character without byte-fiddling.
Parameters:
- `s` — the UTF-8 string
```ludic
program Chars {
entry {
let cs = Unicode.chars("hi")
var i = 0
while i < len(cs) { print(cs[i]); i = i + 1 } # 104, 105
}
}
```

View file

@ -0,0 +1,25 @@
---
id: unicode-grapheme_len
name: Unicode.grapheme_len
category: unicode
kind: namespace-method
tokens: Unicode.grapheme_len
sig: Unicode.grapheme_len(s) -> int
tip: Number of user-perceived characters.
order: 9
ns: Unicode
member: grapheme_len
---
Counts <strong>grapheme clusters</strong> — what a reader perceives as one character — which is what you want for cursor movement and visual width. A base letter plus a combining accent, a ZWJ emoji sequence (like a family emoji), and a two–regional-indicator flag each count as one, even though they span several code points. This v1 approximates UAX#29 over those common cases; full segmentation is a follow-up.
Parameters:
- `s` — the UTF-8 string
```ludic
program GraphemeLen {
entry {
print(Unicode.grapheme_len("hello")) # 5 (plain ASCII)
}
}
```

View file

@ -0,0 +1,25 @@
---
id: unicode-is_valid_utf8
name: Unicode.is_valid_utf8
category: unicode
kind: namespace-method
tokens: Unicode.is_valid_utf8
sig: Unicode.is_valid_utf8(s) -> bool
tip: Is a string well-formed UTF-8?
order: 3
ns: Unicode
member: is_valid_utf8
---
Strictly validates that the bytes are well-formed UTF-8: correct continuation bytes, no overlong encodings, no surrogate code points, and nothing above U+10FFFF. Use it as a gate on untrusted input — a chat line, a downloaded file, a mod's data — before treating it as text.
Parameters:
- `s` — the bytes to validate
```ludic
program Valid {
entry {
if Unicode.is_valid_utf8("safe text") { print(1) } else { print(0) }
}
}
```

View file

@ -0,0 +1,25 @@
---
id: unicode-len
name: Unicode.len
category: unicode
kind: namespace-method
tokens: Unicode.len
sig: Unicode.len(s) -> int
tip: Number of code points in a string (not bytes).
order: 1
ns: Unicode
member: len
---
Counts the <strong>code points</strong> in a UTF-8 string — the correct "length" for most text logic, unlike a raw byte count which over-counts anything non-ASCII. Contrast <a href="unicode-byte_len"><code>byte_len</code></a> (storage size) and <a href="unicode-grapheme_len"><code>grapheme_len</code></a> (user-perceived characters).
Parameters:
- `s` — the UTF-8 string
```ludic
program Len {
entry {
print(Unicode.len("hello")) # 5
}
}
```

View file

@ -0,0 +1,25 @@
---
id: unicode-lower
name: Unicode.lower
category: unicode
kind: namespace-method
tokens: Unicode.lower
sig: Unicode.lower(s) -> string
tip: Lowercase a string.
order: 7
ns: Unicode
member: lower
---
Returns a new string with each letter lowercased. v1 maps ASCII and Latin-1 letters; other scripts and locale-specific rules are follow-ups, and unmapped code points pass through unchanged. Useful for case-insensitive comparison of names and commands.
Parameters:
- `s` — the UTF-8 string
```ludic
program Lower {
entry {
if Unicode.lower("HELLO") == "hello" { print(1) }
}
}
```

View file

@ -0,0 +1,26 @@
---
id: unicode-truncate
name: Unicode.truncate
category: unicode
kind: namespace-method
tokens: Unicode.truncate
sig: Unicode.truncate(s, n) -> string
tip: First n code points of a string.
order: 8
ns: Unicode
member: truncate
---
Returns the first <code>n</code> code points as a new string, never splitting a multibyte character — the safe way to cap a display name or a chat line to a length. If the string has fewer than <code>n</code> code points it is returned whole; <code>n &lt;= 0</code> yields the empty string.
Parameters:
- `s` — the UTF-8 string
- `n` — the maximum number of code points to keep
```ludic
program Truncate {
entry {
if Unicode.truncate("hello", 3) == "hel" { print(1) }
}
}
```

View file

@ -0,0 +1,25 @@
---
id: unicode-upper
name: Unicode.upper
category: unicode
kind: namespace-method
tokens: Unicode.upper
sig: Unicode.upper(s) -> string
tip: Uppercase a string.
order: 6
ns: Unicode
member: upper
---
Returns a new string with each letter uppercased. v1 maps ASCII and Latin-1 letters (correct for Western-European text); other scripts and locale-specific rules are follow-ups, and unmapped code points pass through unchanged.
Parameters:
- `s` — the UTF-8 string
```ludic
program Upper {
entry {
if Unicode.upper("hello") == "HELLO" { print(1) }
}
}
```