All checks were successful
docs / build-and-deploy (push) Successful in 2s
Make Ludic text correct-by-default over UTF-8, so player names, translated UI, and chat behave for every language instead of counting bytes and splitting characters in half. The byte-oriented Text.* stays for speed; Unicode.* is the layer that understands code points and (approximately) grapheme clusters. - len / byte_len code points vs bytes — the two lengths, kept distinct - is_valid_utf8 strict validation of untrusted input - char_at / chars code-point access by index; chars() -> []int - upper / lower case mapping (ASCII + Latin-1) - truncate first n code points, never a half-character - grapheme_len user-perceived characters (approx UAX#29) Pure integer/byte IR over NUL-terminated buffers; C-free, no data-table blob. Decoding and validation cover the full UTF-8 range (overlong/surrogate/>10FFFF rejected). grapheme_len collapses combining marks, variation selectors, ZWJ sequences (family emoji), and regional-indicator flag pairs. Documented v1 scope: wider-script/locale case rules (Latin-Extended, Greek, Cyrillic, Turkish i, German ß) and NFC normalization are follow-ups. - examples/library/unicode.ludic: asserts the invariants across ASCII, Latin-1 (é round-trips through upper/lower), a decomposed "café" (5 code points, 4 graphemes), a ZWJ family emoji (5 code points, 1 grapheme), and a flag (2 regional indicators, 1 grapheme). Wired into `x test` (now 55 passed). - docs: a new Unicode section + 9 per-symbol pages clarifying byte vs code point vs grapheme; inventory updated; every fence passes check-docs; site builds. - seed regenerated; `x bootstrap-cfree` fixpoint holds. Closes #13 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
25 lines
675 B
Markdown
25 lines
675 B
Markdown
---
|
|
id: unicode-len
|
|
name: Unicode.len
|
|
category: unicode
|
|
kind: namespace-method
|
|
tokens: Unicode.len
|
|
sig: Unicode.len(s) -> int
|
|
tip: Number of code points in a string (not bytes).
|
|
order: 1
|
|
ns: Unicode
|
|
member: len
|
|
---
|
|
|
|
Counts the <strong>code points</strong> in a UTF-8 string — the correct "length" for most text logic, unlike a raw byte count which over-counts anything non-ASCII. Contrast <a href="unicode-byte_len"><code>byte_len</code></a> (storage size) and <a href="unicode-grapheme_len"><code>grapheme_len</code></a> (user-perceived characters).
|
|
|
|
Parameters:
|
|
- `s` — the UTF-8 string
|
|
|
|
```ludic
|
|
program Len {
|
|
entry {
|
|
print(Unicode.len("hello")) # 5
|
|
}
|
|
}
|
|
```
|