feat(stdlib): add Unicode.* — UTF-8 code points, graphemes, case mapping (#13)
All checks were successful
docs / build-and-deploy (push) Successful in 2s
All checks were successful
docs / build-and-deploy (push) Successful in 2s
Make Ludic text correct-by-default over UTF-8, so player names, translated UI, and chat behave for every language instead of counting bytes and splitting characters in half. The byte-oriented Text.* stays for speed; Unicode.* is the layer that understands code points and (approximately) grapheme clusters. - len / byte_len code points vs bytes — the two lengths, kept distinct - is_valid_utf8 strict validation of untrusted input - char_at / chars code-point access by index; chars() -> []int - upper / lower case mapping (ASCII + Latin-1) - truncate first n code points, never a half-character - grapheme_len user-perceived characters (approx UAX#29) Pure integer/byte IR over NUL-terminated buffers; C-free, no data-table blob. Decoding and validation cover the full UTF-8 range (overlong/surrogate/>10FFFF rejected). grapheme_len collapses combining marks, variation selectors, ZWJ sequences (family emoji), and regional-indicator flag pairs. Documented v1 scope: wider-script/locale case rules (Latin-Extended, Greek, Cyrillic, Turkish i, German ß) and NFC normalization are follow-ups. - examples/library/unicode.ludic: asserts the invariants across ASCII, Latin-1 (é round-trips through upper/lower), a decomposed "café" (5 code points, 4 graphemes), a ZWJ family emoji (5 code points, 1 grapheme), and a flag (2 regional indicators, 1 grapheme). Wired into `x test` (now 55 passed). - docs: a new Unicode section + 9 per-symbol pages clarifying byte vs code point vs grapheme; inventory updated; every fence passes check-docs; site builds. - seed regenerated; `x bootstrap-cfree` fixpoint holds. Closes #13 Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
b205dd8dfd
commit
bad6cd1ac9
19 changed files with 11934 additions and 10468 deletions
13
docs/language/unicode/_section.md
Normal file
13
docs/language/unicode/_section.md
Normal file
|
|
@ -0,0 +1,13 @@
|
|||
---
|
||||
id: unicode
|
||||
title: Unicode
|
||||
order: 6
|
||||
---
|
||||
|
||||
Correct-by-default text over UTF-8. Player names, translated menus, and chat arrive as UTF-8 bytes, and counting *bytes* gets non-ASCII text wrong — the wrong length, and truncation that slices a character in half into mojibake. The <a href="ns-Unicode"><code>Unicode</code></a> namespace works in <strong>code points</strong> and (approximately) <strong>grapheme clusters</strong> instead, so measuring, indexing, truncating, and case-mapping behave for every language.
|
||||
|
||||
This is the correctness layer, not a replacement: the byte-oriented <a href="ns-Text"><code>Text</code></a> operations stay for speed on ASCII and for raw byte work. Reach for <code>Unicode.*</code> whenever the text came from a human — a name, a message, a localized string.
|
||||
|
||||
Three notions of "length" matter, and the API keeps them distinct: <a href="unicode-byte_len"><code>byte_len</code></a> (storage), <a href="unicode-len"><code>len</code></a> (code points — Unicode scalar values), and <a href="unicode-grapheme_len"><code>grapheme_len</code></a> (user-perceived characters, where a base letter plus its combining accent, or a ZWJ emoji sequence, count as one).
|
||||
|
||||
**Coverage (v1):** decoding and <a href="unicode-is_valid_utf8"><code>validation</code></a> cover the full UTF-8 range. Case mapping covers ASCII and the Latin-1 letters — correct for Western-European text; wider scripts (Latin-Extended, Greek, Cyrillic), locale rules (Turkish <code>i</code>, German <code>ß</code>), and NFC normalization are follow-ups. <a href="unicode-grapheme_len"><code>grapheme_len</code></a> approximates UAX#29 for the cases real player text hits — combining marks, variation selectors, ZWJ sequences, and flag pairs.
|
||||
25
docs/language/unicode/unicode-byte_len.md
Normal file
25
docs/language/unicode/unicode-byte_len.md
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
---
|
||||
id: unicode-byte_len
|
||||
name: Unicode.byte_len
|
||||
category: unicode
|
||||
kind: namespace-method
|
||||
tokens: Unicode.byte_len
|
||||
sig: Unicode.byte_len(s) -> int
|
||||
tip: Number of bytes in a string.
|
||||
order: 2
|
||||
ns: Unicode
|
||||
member: byte_len
|
||||
---
|
||||
|
||||
The number of <strong>bytes</strong> the string occupies — its storage size. This is what raw byte APIs and buffers care about, and it is greater than <a href="unicode-len"><code>len</code></a> whenever the text contains non-ASCII code points (each of which encodes to 2–4 bytes).
|
||||
|
||||
Parameters:
|
||||
- `s` — the UTF-8 string
|
||||
|
||||
```ludic
|
||||
program ByteLen {
|
||||
entry {
|
||||
print(Unicode.byte_len("hello")) # 5 (all ASCII)
|
||||
}
|
||||
}
|
||||
```
|
||||
26
docs/language/unicode/unicode-char_at.md
Normal file
26
docs/language/unicode/unicode-char_at.md
Normal file
|
|
@ -0,0 +1,26 @@
|
|||
---
|
||||
id: unicode-char_at
|
||||
name: Unicode.char_at
|
||||
category: unicode
|
||||
kind: namespace-method
|
||||
tokens: Unicode.char_at
|
||||
sig: Unicode.char_at(s, i) -> int
|
||||
tip: The i-th code point of a string.
|
||||
order: 4
|
||||
ns: Unicode
|
||||
member: char_at
|
||||
---
|
||||
|
||||
Returns the code point at code-point index <code>i</code> as an integer (a Unicode scalar value), or <code>-1</code> when <code>i</code> is out of range. Indexing is by code point, not byte, so it never lands in the middle of a multibyte character.
|
||||
|
||||
Parameters:
|
||||
- `s` — the UTF-8 string
|
||||
- `i` — the code-point index (0-based)
|
||||
|
||||
```ludic
|
||||
program CharAt {
|
||||
entry {
|
||||
print(Unicode.char_at("hello", 0)) # 104 = 'h'
|
||||
}
|
||||
}
|
||||
```
|
||||
27
docs/language/unicode/unicode-chars.md
Normal file
27
docs/language/unicode/unicode-chars.md
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
---
|
||||
id: unicode-chars
|
||||
name: Unicode.chars
|
||||
category: unicode
|
||||
kind: namespace-method
|
||||
tokens: Unicode.chars
|
||||
sig: Unicode.chars(s) -> []int
|
||||
tip: Every code point of a string, in order.
|
||||
order: 5
|
||||
ns: Unicode
|
||||
member: chars
|
||||
---
|
||||
|
||||
Decodes the whole string into a <a href="type-slice"><code>[]int</code></a> of code points, in order — the iteration primitive for laying out glyphs, scanning text, or transforming character by character without byte-fiddling.
|
||||
|
||||
Parameters:
|
||||
- `s` — the UTF-8 string
|
||||
|
||||
```ludic
|
||||
program Chars {
|
||||
entry {
|
||||
let cs = Unicode.chars("hi")
|
||||
var i = 0
|
||||
while i < len(cs) { print(cs[i]); i = i + 1 } # 104, 105
|
||||
}
|
||||
}
|
||||
```
|
||||
25
docs/language/unicode/unicode-grapheme_len.md
Normal file
25
docs/language/unicode/unicode-grapheme_len.md
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
---
|
||||
id: unicode-grapheme_len
|
||||
name: Unicode.grapheme_len
|
||||
category: unicode
|
||||
kind: namespace-method
|
||||
tokens: Unicode.grapheme_len
|
||||
sig: Unicode.grapheme_len(s) -> int
|
||||
tip: Number of user-perceived characters.
|
||||
order: 9
|
||||
ns: Unicode
|
||||
member: grapheme_len
|
||||
---
|
||||
|
||||
Counts <strong>grapheme clusters</strong> — what a reader perceives as one character — which is what you want for cursor movement and visual width. A base letter plus a combining accent, a ZWJ emoji sequence (like a family emoji), and a two–regional-indicator flag each count as one, even though they span several code points. This v1 approximates UAX#29 over those common cases; full segmentation is a follow-up.
|
||||
|
||||
Parameters:
|
||||
- `s` — the UTF-8 string
|
||||
|
||||
```ludic
|
||||
program GraphemeLen {
|
||||
entry {
|
||||
print(Unicode.grapheme_len("hello")) # 5 (plain ASCII)
|
||||
}
|
||||
}
|
||||
```
|
||||
25
docs/language/unicode/unicode-is_valid_utf8.md
Normal file
25
docs/language/unicode/unicode-is_valid_utf8.md
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
---
|
||||
id: unicode-is_valid_utf8
|
||||
name: Unicode.is_valid_utf8
|
||||
category: unicode
|
||||
kind: namespace-method
|
||||
tokens: Unicode.is_valid_utf8
|
||||
sig: Unicode.is_valid_utf8(s) -> bool
|
||||
tip: Is a string well-formed UTF-8?
|
||||
order: 3
|
||||
ns: Unicode
|
||||
member: is_valid_utf8
|
||||
---
|
||||
|
||||
Strictly validates that the bytes are well-formed UTF-8: correct continuation bytes, no overlong encodings, no surrogate code points, and nothing above U+10FFFF. Use it as a gate on untrusted input — a chat line, a downloaded file, a mod's data — before treating it as text.
|
||||
|
||||
Parameters:
|
||||
- `s` — the bytes to validate
|
||||
|
||||
```ludic
|
||||
program Valid {
|
||||
entry {
|
||||
if Unicode.is_valid_utf8("safe text") { print(1) } else { print(0) }
|
||||
}
|
||||
}
|
||||
```
|
||||
25
docs/language/unicode/unicode-len.md
Normal file
25
docs/language/unicode/unicode-len.md
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
---
|
||||
id: unicode-len
|
||||
name: Unicode.len
|
||||
category: unicode
|
||||
kind: namespace-method
|
||||
tokens: Unicode.len
|
||||
sig: Unicode.len(s) -> int
|
||||
tip: Number of code points in a string (not bytes).
|
||||
order: 1
|
||||
ns: Unicode
|
||||
member: len
|
||||
---
|
||||
|
||||
Counts the <strong>code points</strong> in a UTF-8 string — the correct "length" for most text logic, unlike a raw byte count which over-counts anything non-ASCII. Contrast <a href="unicode-byte_len"><code>byte_len</code></a> (storage size) and <a href="unicode-grapheme_len"><code>grapheme_len</code></a> (user-perceived characters).
|
||||
|
||||
Parameters:
|
||||
- `s` — the UTF-8 string
|
||||
|
||||
```ludic
|
||||
program Len {
|
||||
entry {
|
||||
print(Unicode.len("hello")) # 5
|
||||
}
|
||||
}
|
||||
```
|
||||
25
docs/language/unicode/unicode-lower.md
Normal file
25
docs/language/unicode/unicode-lower.md
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
---
|
||||
id: unicode-lower
|
||||
name: Unicode.lower
|
||||
category: unicode
|
||||
kind: namespace-method
|
||||
tokens: Unicode.lower
|
||||
sig: Unicode.lower(s) -> string
|
||||
tip: Lowercase a string.
|
||||
order: 7
|
||||
ns: Unicode
|
||||
member: lower
|
||||
---
|
||||
|
||||
Returns a new string with each letter lowercased. v1 maps ASCII and Latin-1 letters; other scripts and locale-specific rules are follow-ups, and unmapped code points pass through unchanged. Useful for case-insensitive comparison of names and commands.
|
||||
|
||||
Parameters:
|
||||
- `s` — the UTF-8 string
|
||||
|
||||
```ludic
|
||||
program Lower {
|
||||
entry {
|
||||
if Unicode.lower("HELLO") == "hello" { print(1) }
|
||||
}
|
||||
}
|
||||
```
|
||||
26
docs/language/unicode/unicode-truncate.md
Normal file
26
docs/language/unicode/unicode-truncate.md
Normal file
|
|
@ -0,0 +1,26 @@
|
|||
---
|
||||
id: unicode-truncate
|
||||
name: Unicode.truncate
|
||||
category: unicode
|
||||
kind: namespace-method
|
||||
tokens: Unicode.truncate
|
||||
sig: Unicode.truncate(s, n) -> string
|
||||
tip: First n code points of a string.
|
||||
order: 8
|
||||
ns: Unicode
|
||||
member: truncate
|
||||
---
|
||||
|
||||
Returns the first <code>n</code> code points as a new string, never splitting a multibyte character — the safe way to cap a display name or a chat line to a length. If the string has fewer than <code>n</code> code points it is returned whole; <code>n <= 0</code> yields the empty string.
|
||||
|
||||
Parameters:
|
||||
- `s` — the UTF-8 string
|
||||
- `n` — the maximum number of code points to keep
|
||||
|
||||
```ludic
|
||||
program Truncate {
|
||||
entry {
|
||||
if Unicode.truncate("hello", 3) == "hel" { print(1) }
|
||||
}
|
||||
}
|
||||
```
|
||||
25
docs/language/unicode/unicode-upper.md
Normal file
25
docs/language/unicode/unicode-upper.md
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
---
|
||||
id: unicode-upper
|
||||
name: Unicode.upper
|
||||
category: unicode
|
||||
kind: namespace-method
|
||||
tokens: Unicode.upper
|
||||
sig: Unicode.upper(s) -> string
|
||||
tip: Uppercase a string.
|
||||
order: 6
|
||||
ns: Unicode
|
||||
member: upper
|
||||
---
|
||||
|
||||
Returns a new string with each letter uppercased. v1 maps ASCII and Latin-1 letters (correct for Western-European text); other scripts and locale-specific rules are follow-ups, and unmapped code points pass through unchanged.
|
||||
|
||||
Parameters:
|
||||
- `s` — the UTF-8 string
|
||||
|
||||
```ludic
|
||||
program Upper {
|
||||
entry {
|
||||
if Unicode.upper("hello") == "HELLO" { print(1) }
|
||||
}
|
||||
}
|
||||
```
|
||||
49
examples/library/unicode.ludic
Normal file
49
examples/library/unicode.ludic
Normal file
|
|
@ -0,0 +1,49 @@
|
|||
# unicode.ludic — Unicode.* correctness over UTF-8. Byte counting is wrong for
|
||||
# non-ASCII, so we assert the code-point / grapheme invariants that must hold:
|
||||
# len counts code points not bytes, validation accepts good UTF-8, char access
|
||||
# never splits a character, case mapping round-trips, truncate keeps whole
|
||||
# characters, and grapheme_len collapses combining marks, ZWJ emoji, and flags.
|
||||
# Running it prints: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18
|
||||
program Unicode {
|
||||
entry {
|
||||
# code points vs bytes
|
||||
if Unicode.len("héllo") == 5 { print(1) }
|
||||
if Unicode.byte_len("héllo") == 6 { print(2) }
|
||||
if Unicode.len("abc") == 3 { print(3) }
|
||||
|
||||
# validation of well-formed UTF-8
|
||||
if Unicode.is_valid_utf8("héllo") { print(4) }
|
||||
|
||||
# code-point access by index (é is U+00E9 = 233), out of range -> -1
|
||||
if Unicode.char_at("héllo", 0) == 104 { print(5) }
|
||||
if Unicode.char_at("héllo", 1) == 233 { print(6) }
|
||||
if Unicode.char_at("héllo", 5) < 0 { print(7) }
|
||||
|
||||
# chars() yields every code point in order
|
||||
let cs = Unicode.chars("héllo")
|
||||
if len(cs) == 5 { print(8) }
|
||||
if cs[1] == 233 { print(9) }
|
||||
|
||||
# case mapping (ASCII + Latin-1) round-trips
|
||||
if Unicode.upper("héllo") == "HÉLLO" { print(10) }
|
||||
if Unicode.lower("HÉLLO") == "héllo" { print(11) }
|
||||
|
||||
# truncate keeps whole characters: first 3 code points of "héllo" is "hél"
|
||||
if Unicode.truncate("héllo", 3) == "hél" { print(12) }
|
||||
|
||||
# grapheme clustering: ASCII is one grapheme per code point
|
||||
if Unicode.grapheme_len("abc") == 3 { print(13) }
|
||||
|
||||
# a combining accent joins its base: "café" decomposed is 5 code points but
|
||||
# 4 user-perceived characters
|
||||
if Unicode.grapheme_len("café") == 4 { print(14) }
|
||||
if Unicode.len("café") == 5 { print(15) }
|
||||
|
||||
# a ZWJ emoji sequence (family) is 5 code points but one grapheme
|
||||
if Unicode.grapheme_len("👨👩👧") == 1 { print(16) }
|
||||
if Unicode.len("👨👩👧") == 5 { print(17) }
|
||||
|
||||
# a flag is two regional indicators but one grapheme
|
||||
if Unicode.grapheme_len("🇺🇸") == 1 { print(18) }
|
||||
}
|
||||
}
|
||||
|
|
@ -30,6 +30,7 @@ var g_uses_uuidrt: bool = false # Uuid.* was emitted -> emit the UUID runtime (
|
|||
var g_uses_noisert: bool = false # Noise.* was emitted -> emit the fixed-point noise runtime
|
||||
var g_uses_logrt: bool = false # Log.* was emitted -> emit the log level register + console sink
|
||||
var g_uses_osrt: bool = false # Os.* (prelude-backed methods) was emitted -> emit the Os runtime
|
||||
var g_uses_unicodert: bool = false # Unicode.* was emitted -> emit the UTF-8 runtime
|
||||
var g_uses_datert: bool = false # Date.*/DateTime.* was emitted -> emit the civil<->epoch conversions
|
||||
var g_uses_longstr: bool = false # string(long) / interpolating a long was emitted -> emit fn_long_str
|
||||
|
||||
|
|
|
|||
|
|
@ -109,6 +109,7 @@ function emit_program() -> void {
|
|||
if g_uses_noisert { emit_noise_prelude() } # @fn_noise_value2/perlin2/simplex2/fbm2/cellular2 (Q16.16)
|
||||
if g_uses_logrt { emit_log_prelude() } # @L_log_level + @fn_log_emit (levelled stderr sink)
|
||||
if g_uses_osrt { emit_os_prelude() } # @fn_os_args/platform/arch/save_dir/... (libc env + uname)
|
||||
if g_uses_unicodert { emit_unicode_prelude() } # @fn_uni_len/valid/decode/case/truncate/grapheme (UTF-8)
|
||||
if g_uses_datert { emit_datetime_prelude() } # @fn_days_from_civil / @fn_civil_from_days conversions
|
||||
}
|
||||
|
||||
|
|
|
|||
|
|
@ -251,6 +251,10 @@ function emit_ns_call(ns: pointer, meth: pointer, e: Node) -> Val {
|
|||
if is_os_ns(meth) { return emit_os_ns(meth, e) }
|
||||
perr(`unknown builtin Os.{meth}`)
|
||||
}
|
||||
if (ns == "Unicode") {
|
||||
if is_unicode_ns(meth) { return emit_unicode_ns(meth, e) }
|
||||
perr(`unknown builtin Unicode.{meth}`)
|
||||
}
|
||||
if (ns == "Vector") {
|
||||
if is_vector_ns(meth) { return emit_vector_ns(meth, e) }
|
||||
perr(`unknown builtin Vector.{meth}`)
|
||||
|
|
|
|||
318
selfhost/emit_unicode.ludic
Normal file
318
selfhost/emit_unicode.ludic
Normal file
|
|
@ -0,0 +1,318 @@
|
|||
# emit_unicode.ludic — the Unicode.* namespace: correct-by-default text over
|
||||
# UTF-8, so player names, translated UI, and chat behave for every language
|
||||
# instead of counting bytes and splitting characters in half. The byte-oriented
|
||||
# Text.* stays for speed; Unicode.* is the layer that understands code points and
|
||||
# (approximately) grapheme clusters.
|
||||
#
|
||||
# Unicode.len(s) -> int number of code points (not bytes)
|
||||
# Unicode.byte_len(s) -> int number of bytes (the contrast to len)
|
||||
# Unicode.is_valid_utf8(s) -> bool strict UTF-8 validation of untrusted input
|
||||
# Unicode.char_at(s, i) -> int the i-th code point (-1 if out of range)
|
||||
# Unicode.chars(s) -> []int every code point, in order
|
||||
# Unicode.upper(s) -> string uppercased (ASCII + Latin-1)
|
||||
# Unicode.lower(s) -> string lowercased (ASCII + Latin-1)
|
||||
# Unicode.truncate(s, n) -> string first n code points, never a half-char
|
||||
# Unicode.grapheme_len(s) -> int user-perceived characters (approx UAX#29)
|
||||
#
|
||||
# Unicode version / coverage: this v1 decodes and validates the full UTF-8 range.
|
||||
# Case mapping covers ASCII and the Latin-1 Supplement letters (correct for
|
||||
# Western-European text); Latin-Extended/Greek/Cyrillic case, locale rules (e.g.
|
||||
# Turkish i, German ß->SS), and normalization (NFC) are documented follow-ups.
|
||||
# grapheme_len is an approximation of UAX#29 that handles combining marks,
|
||||
# variation selectors, ZWJ sequences (e.g. family emoji), and regional-indicator
|
||||
# flag pairs — enough for measuring and truncating real player text; full
|
||||
# segmentation with its data tables is a follow-up.
|
||||
#
|
||||
# Every function is pure over a NUL-terminated UTF-8 buffer: no global state, no
|
||||
# platform dependence, so results are identical on native, headless, and wasm.
|
||||
|
||||
function is_unicode_ns(meth: pointer) -> bool {
|
||||
if (meth == "len") or (meth == "byte_len") or (meth == "is_valid_utf8") { return true }
|
||||
if (meth == "char_at") or (meth == "chars") { return true }
|
||||
if (meth == "upper") or (meth == "lower") or (meth == "truncate") { return true }
|
||||
if (meth == "grapheme_len") { return true }
|
||||
return false
|
||||
}
|
||||
|
||||
function emit_unicode_ns(meth: pointer, e: Node) -> Val {
|
||||
g_uses_unicodert = true
|
||||
if (meth == "byte_len") { # bytes, the contrast to len()
|
||||
let s = emit_expr(e.kids[0])
|
||||
let r = emit_bind(`call i64 @strlen(ptr {s.code})`)
|
||||
return val(emit_bind(`trunc i64 {r} to i32`), "int")
|
||||
}
|
||||
if (meth == "len") {
|
||||
let s = emit_expr(e.kids[0])
|
||||
return val(emit_bind(`call i32 @fn_uni_len(ptr {s.code})`), "int")
|
||||
}
|
||||
if (meth == "is_valid_utf8") {
|
||||
let s = emit_expr(e.kids[0])
|
||||
return val(emit_bind(`call i32 @fn_uni_valid(ptr {s.code})`), "bool")
|
||||
}
|
||||
if (meth == "char_at") {
|
||||
let s = emit_expr(e.kids[0]); let i = emit_expr(e.kids[1])
|
||||
return val(emit_bind(`call i32 @fn_uni_char_at(ptr {s.code}, i32 {i.code})`), "int")
|
||||
}
|
||||
if (meth == "chars") {
|
||||
let s = emit_expr(e.kids[0])
|
||||
return val(emit_bind(`call ptr @fn_uni_chars(ptr {s.code})`), "[]int")
|
||||
}
|
||||
if (meth == "upper") {
|
||||
let s = emit_expr(e.kids[0])
|
||||
return val(emit_bind(`call ptr @fn_uni_case(ptr {s.code}, i32 1)`), "string")
|
||||
}
|
||||
if (meth == "lower") {
|
||||
let s = emit_expr(e.kids[0])
|
||||
return val(emit_bind(`call ptr @fn_uni_case(ptr {s.code}, i32 0)`), "string")
|
||||
}
|
||||
if (meth == "truncate") {
|
||||
let s = emit_expr(e.kids[0]); let n = emit_expr(e.kids[1])
|
||||
return val(emit_bind(`call ptr @fn_uni_truncate(ptr {s.code}, i32 {n.code})`), "string")
|
||||
}
|
||||
# grapheme_len
|
||||
let s = emit_expr(e.kids[0])
|
||||
return val(emit_bind(`call i32 @fn_uni_grapheme_len(ptr {s.code})`), "int")
|
||||
}
|
||||
|
||||
# emit_unicode_prelude — the UTF-8 runtime, emitted once per program that uses
|
||||
# Unicode.* (g_uses_unicodert). Pure integer/byte IR over NUL-terminated buffers.
|
||||
function emit_unicode_prelude() -> void {
|
||||
emit_uni_core()
|
||||
emit_uni_case_fns()
|
||||
emit_uni_str_fns()
|
||||
}
|
||||
|
||||
# ---- decode + length + validation ------------------------------------------
|
||||
function emit_uni_core() -> void {
|
||||
# code-point count: every byte that is NOT a UTF-8 continuation byte
|
||||
# (0b10xxxxxx) begins a new code point. Walks byte-by-byte, so it is safe on
|
||||
# truncated/invalid input and stops exactly at the NUL.
|
||||
emith("define i32 @fn_uni_len(ptr %s) {\n")
|
||||
emith("entry:\n %ip = alloca i32\n %np = alloca i32\n store i32 0, ptr %ip\n store i32 0, ptr %np\n br label %lp\n")
|
||||
emith("lp:\n %i = load i32, ptr %ip\n %p = getelementptr i8, ptr %s, i32 %i\n %b = load i8, ptr %p\n %c = zext i8 %b to i32\n")
|
||||
emith(" %z = icmp eq i32 %c, 0\n br i1 %z, label %done, label %go\n")
|
||||
emith("go:\n %m = and i32 %c, 192\n %cont = icmp eq i32 %m, 128\n br i1 %cont, label %skip, label %count\n")
|
||||
emith("count:\n %n = load i32, ptr %np\n %n1 = add i32 %n, 1\n store i32 %n1, ptr %np\n br label %skip\n")
|
||||
emith("skip:\n %i1 = add i32 %i, 1\n store i32 %i1, ptr %ip\n br label %lp\n")
|
||||
emith("done:\n %r = load i32, ptr %np\n ret i32 %r\n}\n")
|
||||
|
||||
# decode one code point at byte offset %i; store it to %cpout; return the byte
|
||||
# offset just past it. Lenient and overrun-safe: a truncated multibyte sequence
|
||||
# (a continuation byte that is NUL) or an invalid lead byte decodes as a single
|
||||
# byte, so the walk always makes progress and never reads past the terminator.
|
||||
emith("define i32 @fn_uni_decode(ptr %s, i32 %i, ptr %cpout) {\n")
|
||||
emith("entry:\n %p0 = getelementptr i8, ptr %s, i32 %i\n %b0 = load i8, ptr %p0\n %c0 = zext i8 %b0 to i32\n")
|
||||
emith(" %a1 = icmp ult i32 %c0, 128\n br i1 %a1, label %one, label %multi\n")
|
||||
emith("one:\n store i32 %c0, ptr %cpout\n %oi = add i32 %i, 1\n ret i32 %oi\n")
|
||||
emith("multi:\n %h2 = and i32 %c0, 224\n %is2 = icmp eq i32 %h2, 192\n br i1 %is2, label %two, label %c3\n")
|
||||
emith("c3:\n %h3 = and i32 %c0, 240\n %is3 = icmp eq i32 %h3, 224\n br i1 %is3, label %three, label %c4\n")
|
||||
emith("c4:\n %h4 = and i32 %c0, 248\n %is4 = icmp eq i32 %h4, 240\n br i1 %is4, label %four, label %bad\n")
|
||||
emith("bad:\n store i32 %c0, ptr %cpout\n %bi = add i32 %i, 1\n ret i32 %bi\n")
|
||||
# two-byte
|
||||
emith("two:\n %t1i = add i32 %i, 1\n %t1p = getelementptr i8, ptr %s, i32 %t1i\n %t1b = load i8, ptr %t1p\n %t1 = zext i8 %t1b to i32\n")
|
||||
emith(" %t1z = icmp eq i32 %t1, 0\n br i1 %t1z, label %tbad, label %tok\n")
|
||||
emith("tbad:\n store i32 %c0, ptr %cpout\n %tbi = add i32 %i, 1\n ret i32 %tbi\n")
|
||||
emith("tok:\n %tm = and i32 %c0, 31\n %tsh = shl i32 %tm, 6\n %tl = and i32 %t1, 63\n %tcp = or i32 %tsh, %tl\n store i32 %tcp, ptr %cpout\n %ti = add i32 %i, 2\n ret i32 %ti\n")
|
||||
# three-byte
|
||||
emith("three:\n %r1i = add i32 %i, 1\n %r1p = getelementptr i8, ptr %s, i32 %r1i\n %r1b = load i8, ptr %r1p\n %r1 = zext i8 %r1b to i32\n")
|
||||
emith(" %r1z = icmp eq i32 %r1, 0\n br i1 %r1z, label %rbad, label %rc2\n")
|
||||
emith("rbad:\n store i32 %c0, ptr %cpout\n %rbi = add i32 %i, 1\n ret i32 %rbi\n")
|
||||
emith("rc2:\n %r2i = add i32 %i, 2\n %r2p = getelementptr i8, ptr %s, i32 %r2i\n %r2b = load i8, ptr %r2p\n %r2 = zext i8 %r2b to i32\n")
|
||||
emith(" %r2z = icmp eq i32 %r2, 0\n br i1 %r2z, label %rbad, label %rok\n")
|
||||
emith("rok:\n %rm = and i32 %c0, 15\n %rsh = shl i32 %rm, 12\n %ram = and i32 %r1, 63\n %rash = shl i32 %ram, 6\n %rbm = and i32 %r2, 63\n")
|
||||
emith(" %rp1 = or i32 %rsh, %rash\n %rcp = or i32 %rp1, %rbm\n store i32 %rcp, ptr %cpout\n %ri = add i32 %i, 3\n ret i32 %ri\n")
|
||||
# four-byte
|
||||
emith("four:\n %f1i = add i32 %i, 1\n %f1p = getelementptr i8, ptr %s, i32 %f1i\n %f1b = load i8, ptr %f1p\n %f1 = zext i8 %f1b to i32\n")
|
||||
emith(" %f1z = icmp eq i32 %f1, 0\n br i1 %f1z, label %fbad, label %fc2\n")
|
||||
emith("fbad:\n store i32 %c0, ptr %cpout\n %fbi = add i32 %i, 1\n ret i32 %fbi\n")
|
||||
emith("fc2:\n %f2i = add i32 %i, 2\n %f2p = getelementptr i8, ptr %s, i32 %f2i\n %f2b = load i8, ptr %f2p\n %f2 = zext i8 %f2b to i32\n")
|
||||
emith(" %f2z = icmp eq i32 %f2, 0\n br i1 %f2z, label %fbad, label %fc3\n")
|
||||
emith("fc3:\n %f3i = add i32 %i, 3\n %f3p = getelementptr i8, ptr %s, i32 %f3i\n %f3b = load i8, ptr %f3p\n %f3 = zext i8 %f3b to i32\n")
|
||||
emith(" %f3z = icmp eq i32 %f3, 0\n br i1 %f3z, label %fbad, label %fok\n")
|
||||
emith("fok:\n %fm = and i32 %c0, 7\n %fsh = shl i32 %fm, 18\n %fam = and i32 %f1, 63\n %fash = shl i32 %fam, 12\n")
|
||||
emith(" %fbm = and i32 %f2, 63\n %fbsh = shl i32 %fbm, 6\n %fcm = and i32 %f3, 63\n")
|
||||
emith(" %fp1 = or i32 %fsh, %fash\n %fp2 = or i32 %fp1, %fbsh\n %fcp = or i32 %fp2, %fcm\n store i32 %fcp, ptr %cpout\n %fi = add i32 %i, 4\n ret i32 %fi\n}\n")
|
||||
|
||||
# strict UTF-8 validation: correct continuation bytes, no overlong encodings,
|
||||
# no surrogates (U+D800..U+DFFF), and nothing above U+10FFFF. Returns 1/0.
|
||||
emith("define i32 @fn_uni_valid(ptr %s) {\n")
|
||||
emith("entry:\n %ip = alloca i32\n store i32 0, ptr %ip\n br label %lp\n")
|
||||
emith("lp:\n %i = load i32, ptr %ip\n %p = getelementptr i8, ptr %s, i32 %i\n %b = load i8, ptr %p\n %c = zext i8 %b to i32\n")
|
||||
emith(" %z = icmp eq i32 %c, 0\n br i1 %z, label %good, label %g\n")
|
||||
emith("g:\n %a1 = icmp ult i32 %c, 128\n br i1 %a1, label %adv1, label %m2\n")
|
||||
emith("adv1:\n %n1 = add i32 %i, 1\n store i32 %n1, ptr %ip\n br label %lp\n")
|
||||
emith("m2:\n %h2 = and i32 %c, 224\n %is2 = icmp eq i32 %h2, 192\n br i1 %is2, label %do2, label %m3\n")
|
||||
emith("m3:\n %h3 = and i32 %c, 240\n %is3 = icmp eq i32 %h3, 224\n br i1 %is3, label %do3, label %m4\n")
|
||||
emith("m4:\n %h4 = and i32 %c, 248\n %is4 = icmp eq i32 %h4, 240\n br i1 %is4, label %do4, label %bad\n")
|
||||
emith("bad:\n ret i32 0\n")
|
||||
# 2-byte: cp in [0x80,0x7FF]
|
||||
emith("do2:\n %v1i = add i32 %i, 1\n %v1p = getelementptr i8, ptr %s, i32 %v1i\n %v1b = load i8, ptr %v1p\n %v1 = zext i8 %v1b to i32\n")
|
||||
emith(" %v1m = and i32 %v1, 192\n %v1ok = icmp eq i32 %v1m, 128\n br i1 %v1ok, label %do2b, label %bad\n")
|
||||
emith("do2b:\n %q2m = and i32 %c, 31\n %q2s = shl i32 %q2m, 6\n %q2l = and i32 %v1, 63\n %cp2 = or i32 %q2s, %q2l\n")
|
||||
emith(" %ov2 = icmp ult i32 %cp2, 128\n br i1 %ov2, label %bad, label %adv2\n")
|
||||
emith("adv2:\n %n2 = add i32 %i, 2\n store i32 %n2, ptr %ip\n br label %lp\n")
|
||||
# 3-byte: cp in [0x800,0xFFFF], excluding surrogates
|
||||
emith("do3:\n %w1i = add i32 %i, 1\n %w1p = getelementptr i8, ptr %s, i32 %w1i\n %w1b = load i8, ptr %w1p\n %w1 = zext i8 %w1b to i32\n")
|
||||
emith(" %w1m = and i32 %w1, 192\n %w1ok = icmp eq i32 %w1m, 128\n br i1 %w1ok, label %do3b, label %bad\n")
|
||||
emith("do3b:\n %w2i = add i32 %i, 2\n %w2p = getelementptr i8, ptr %s, i32 %w2i\n %w2b = load i8, ptr %w2p\n %w2 = zext i8 %w2b to i32\n")
|
||||
emith(" %w2m = and i32 %w2, 192\n %w2ok = icmp eq i32 %w2m, 128\n br i1 %w2ok, label %do3c, label %bad\n")
|
||||
emith("do3c:\n %e3m = and i32 %c, 15\n %e3s = shl i32 %e3m, 12\n %e3am = and i32 %w1, 63\n %e3as = shl i32 %e3am, 6\n %e3bm = and i32 %w2, 63\n")
|
||||
emith(" %cp3p = or i32 %e3s, %e3as\n %cp3 = or i32 %cp3p, %e3bm\n")
|
||||
emith(" %ov3 = icmp ult i32 %cp3, 2048\n br i1 %ov3, label %bad, label %do3d\n")
|
||||
emith("do3d:\n %sg1 = icmp uge i32 %cp3, 55296\n %sg2 = icmp ule i32 %cp3, 57343\n %sg = and i1 %sg1, %sg2\n br i1 %sg, label %bad, label %adv3\n")
|
||||
emith("adv3:\n %n3 = add i32 %i, 3\n store i32 %n3, ptr %ip\n br label %lp\n")
|
||||
# 4-byte: cp in [0x10000,0x10FFFF]
|
||||
emith("do4:\n %x1i = add i32 %i, 1\n %x1p = getelementptr i8, ptr %s, i32 %x1i\n %x1b = load i8, ptr %x1p\n %x1 = zext i8 %x1b to i32\n")
|
||||
emith(" %x1m = and i32 %x1, 192\n %x1ok = icmp eq i32 %x1m, 128\n br i1 %x1ok, label %do4b, label %bad\n")
|
||||
emith("do4b:\n %x2i = add i32 %i, 2\n %x2p = getelementptr i8, ptr %s, i32 %x2i\n %x2b = load i8, ptr %x2p\n %x2 = zext i8 %x2b to i32\n")
|
||||
emith(" %x2m = and i32 %x2, 192\n %x2ok = icmp eq i32 %x2m, 128\n br i1 %x2ok, label %do4c, label %bad\n")
|
||||
emith("do4c:\n %x3i = add i32 %i, 3\n %x3p = getelementptr i8, ptr %s, i32 %x3i\n %x3b = load i8, ptr %x3p\n %x3 = zext i8 %x3b to i32\n")
|
||||
emith(" %x3m = and i32 %x3, 192\n %x3ok = icmp eq i32 %x3m, 128\n br i1 %x3ok, label %do4d, label %bad\n")
|
||||
emith("do4d:\n %y0 = and i32 %c, 7\n %y0s = shl i32 %y0, 18\n %y1m = and i32 %x1, 63\n %y1s = shl i32 %y1m, 12\n %y2m = and i32 %x2, 63\n %y2s = shl i32 %y2m, 6\n %y3m = and i32 %x3, 63\n")
|
||||
emith(" %cp4p = or i32 %y0s, %y1s\n %cp4q = or i32 %cp4p, %y2s\n %cp4 = or i32 %cp4q, %y3m\n")
|
||||
emith(" %ov4 = icmp ult i32 %cp4, 65536\n br i1 %ov4, label %bad, label %do4e\n")
|
||||
emith("do4e:\n %hi4 = icmp ugt i32 %cp4, 1114111\n br i1 %hi4, label %bad, label %adv4\n")
|
||||
emith("adv4:\n %n4 = add i32 %i, 4\n store i32 %n4, ptr %ip\n br label %lp\n")
|
||||
emith("good:\n ret i32 1\n}\n")
|
||||
}
|
||||
|
||||
# ---- case mapping + encode helpers -----------------------------------------
|
||||
function emit_uni_case_fns() -> void {
|
||||
# uppercase one code point: ASCII a-z and Latin-1 a-with-diacritic .. thorn
|
||||
# (0xE0..0xFE except 0xF7). Other code points pass through unchanged (v1).
|
||||
emith("define i32 @fn_uni_upcp(i32 %c) {\n")
|
||||
emith(" %la = icmp uge i32 %c, 97\n %lb = icmp ule i32 %c, 122\n %asc = and i1 %la, %lb\n")
|
||||
emith(" %da = icmp uge i32 %c, 224\n %db = icmp ule i32 %c, 254\n %dd = icmp ne i32 %c, 247\n %d1 = and i1 %da, %db\n %d2 = and i1 %d1, %dd\n")
|
||||
emith(" %map = or i1 %asc, %d2\n %up = sub i32 %c, 32\n %r = select i1 %map, i32 %up, i32 %c\n ret i32 %r\n}\n")
|
||||
|
||||
# lowercase one code point: ASCII A-Z and Latin-1 A-with-diacritic .. Thorn
|
||||
# (0xC0..0xDE except 0xD7). Other code points pass through unchanged (v1).
|
||||
emith("define i32 @fn_uni_locp(i32 %c) {\n")
|
||||
emith(" %ua = icmp uge i32 %c, 65\n %ub = icmp ule i32 %c, 90\n %asc = and i1 %ua, %ub\n")
|
||||
emith(" %da = icmp uge i32 %c, 192\n %db = icmp ule i32 %c, 222\n %dd = icmp ne i32 %c, 215\n %d1 = and i1 %da, %db\n %d2 = and i1 %d1, %dd\n")
|
||||
emith(" %map = or i1 %asc, %d2\n %lo = add i32 %c, 32\n %r = select i1 %map, i32 %lo, i32 %c\n ret i32 %r\n}\n")
|
||||
|
||||
# UTF-8 byte width needed to encode a code point
|
||||
emith("define i32 @fn_uni_cpwidth(i32 %cp) {\n")
|
||||
emith(" %a = icmp ult i32 %cp, 128\n %b = icmp ult i32 %cp, 2048\n %c = icmp ult i32 %cp, 65536\n")
|
||||
emith(" %w34 = select i1 %c, i32 3, i32 4\n %w234 = select i1 %b, i32 2, i32 %w34\n %w = select i1 %a, i32 1, i32 %w234\n ret i32 %w\n}\n")
|
||||
|
||||
# encode a code point into %dst at byte offset %off; return the new offset
|
||||
emith("define i32 @fn_uni_encode(ptr %dst, i32 %off, i32 %cp) {\n")
|
||||
emith("entry:\n %w = call i32 @fn_uni_cpwidth(i32 %cp)\n %is1 = icmp eq i32 %w, 1\n br i1 %is1, label %e1, label %k2\n")
|
||||
emith("e1:\n %d0 = getelementptr i8, ptr %dst, i32 %off\n %b0 = trunc i32 %cp to i8\n store i8 %b0, ptr %d0\n %o1 = add i32 %off, 1\n ret i32 %o1\n")
|
||||
emith("k2:\n %is2 = icmp eq i32 %w, 2\n br i1 %is2, label %e2, label %k3\n")
|
||||
emith("e2:\n %hi2 = lshr i32 %cp, 6\n %by0 = or i32 %hi2, 192\n %lo2 = and i32 %cp, 63\n %by1 = or i32 %lo2, 128\n")
|
||||
emith(" %p20 = getelementptr i8, ptr %dst, i32 %off\n %t20 = trunc i32 %by0 to i8\n store i8 %t20, ptr %p20\n")
|
||||
emith(" %o21 = add i32 %off, 1\n %p21 = getelementptr i8, ptr %dst, i32 %o21\n %t21 = trunc i32 %by1 to i8\n store i8 %t21, ptr %p21\n %o2 = add i32 %off, 2\n ret i32 %o2\n")
|
||||
emith("k3:\n %is3 = icmp eq i32 %w, 3\n br i1 %is3, label %e3, label %e4\n")
|
||||
emith("e3:\n %hi3 = lshr i32 %cp, 12\n %by30 = or i32 %hi3, 224\n %m31 = lshr i32 %cp, 6\n %m31b = and i32 %m31, 63\n %by31 = or i32 %m31b, 128\n %m32 = and i32 %cp, 63\n %by32 = or i32 %m32, 128\n")
|
||||
emith(" %p30 = getelementptr i8, ptr %dst, i32 %off\n %t30 = trunc i32 %by30 to i8\n store i8 %t30, ptr %p30\n")
|
||||
emith(" %o31 = add i32 %off, 1\n %p31 = getelementptr i8, ptr %dst, i32 %o31\n %t31 = trunc i32 %by31 to i8\n store i8 %t31, ptr %p31\n")
|
||||
emith(" %o32 = add i32 %off, 2\n %p32 = getelementptr i8, ptr %dst, i32 %o32\n %t32 = trunc i32 %by32 to i8\n store i8 %t32, ptr %p32\n %o3 = add i32 %off, 3\n ret i32 %o3\n")
|
||||
emith("e4:\n %hi4 = lshr i32 %cp, 18\n %by40 = or i32 %hi4, 240\n %m41 = lshr i32 %cp, 12\n %m41b = and i32 %m41, 63\n %by41 = or i32 %m41b, 128\n")
|
||||
emith(" %m42 = lshr i32 %cp, 6\n %m42b = and i32 %m42, 63\n %by42 = or i32 %m42b, 128\n %m43 = and i32 %cp, 63\n %by43 = or i32 %m43, 128\n")
|
||||
emith(" %p40 = getelementptr i8, ptr %dst, i32 %off\n %t40 = trunc i32 %by40 to i8\n store i8 %t40, ptr %p40\n")
|
||||
emith(" %o41 = add i32 %off, 1\n %p41 = getelementptr i8, ptr %dst, i32 %o41\n %t41 = trunc i32 %by41 to i8\n store i8 %t41, ptr %p41\n")
|
||||
emith(" %o42 = add i32 %off, 2\n %p42 = getelementptr i8, ptr %dst, i32 %o42\n %t42 = trunc i32 %by42 to i8\n store i8 %t42, ptr %p42\n")
|
||||
emith(" %o43 = add i32 %off, 3\n %p43 = getelementptr i8, ptr %dst, i32 %o43\n %t43 = trunc i32 %by43 to i8\n store i8 %t43, ptr %p43\n %o4 = add i32 %off, 4\n ret i32 %o4\n}\n")
|
||||
|
||||
# is this code point a grapheme "extend" (a combining mark or variation
|
||||
# selector that joins the preceding base)? An approximation of the common
|
||||
# ranges; ZWJ and regional indicators are handled by grapheme_len itself.
|
||||
emith("define i32 @fn_uni_is_extend(i32 %c) {\n")
|
||||
emith(" %r1 = call i32 @fn_uni_inrange(i32 %c, i32 768, i32 879)\n") # 0300-036F combining diacritics
|
||||
emith(" %r2 = call i32 @fn_uni_inrange(i32 %c, i32 6832, i32 6911)\n") # 1AB0-1AFF
|
||||
emith(" %r3 = call i32 @fn_uni_inrange(i32 %c, i32 7616, i32 7679)\n") # 1DC0-1DFF
|
||||
emith(" %r4 = call i32 @fn_uni_inrange(i32 %c, i32 8400, i32 8447)\n") # 20D0-20FF
|
||||
emith(" %r5 = call i32 @fn_uni_inrange(i32 %c, i32 65056, i32 65071)\n") # FE20-FE2F
|
||||
emith(" %r6 = call i32 @fn_uni_inrange(i32 %c, i32 65024, i32 65039)\n") # FE00-FE0F variation selectors
|
||||
emith(" %r7 = call i32 @fn_uni_inrange(i32 %c, i32 917760, i32 917999)\n") # E0100-E01EF
|
||||
emith(" %o1 = or i32 %r1, %r2\n %o2 = or i32 %o1, %r3\n %o3 = or i32 %o2, %r4\n %o4 = or i32 %o3, %r5\n %o5 = or i32 %o4, %r6\n %o6 = or i32 %o5, %r7\n ret i32 %o6\n}\n")
|
||||
|
||||
emith("define i32 @fn_uni_inrange(i32 %c, i32 %lo, i32 %hi) {\n")
|
||||
emith(" %a = icmp uge i32 %c, %lo\n %b = icmp ule i32 %c, %hi\n %x = and i1 %a, %b\n %r = zext i1 %x to i32\n ret i32 %r\n}\n")
|
||||
}
|
||||
|
||||
# ---- code-point access + string builders -----------------------------------
|
||||
function emit_uni_str_fns() -> void {
|
||||
# the idx-th code point, or -1 when idx is past the end
|
||||
emith("define i32 @fn_uni_char_at(ptr %s, i32 %idx) {\n")
|
||||
emith("entry:\n %ip = alloca i32\n %kp = alloca i32\n %cp = alloca i32\n store i32 0, ptr %ip\n store i32 0, ptr %kp\n br label %lp\n")
|
||||
emith("lp:\n %i = load i32, ptr %ip\n %p = getelementptr i8, ptr %s, i32 %i\n %b = load i8, ptr %p\n %c = zext i8 %b to i32\n")
|
||||
emith(" %z = icmp eq i32 %c, 0\n br i1 %z, label %none, label %go\n")
|
||||
emith("go:\n %ni = call i32 @fn_uni_decode(ptr %s, i32 %i, ptr %cp)\n %k = load i32, ptr %kp\n %hit = icmp eq i32 %k, %idx\n br i1 %hit, label %found, label %next\n")
|
||||
emith("found:\n %v = load i32, ptr %cp\n ret i32 %v\n")
|
||||
emith("next:\n %k1 = add i32 %k, 1\n store i32 %k1, ptr %kp\n store i32 %ni, ptr %ip\n br label %lp\n")
|
||||
emith("none:\n ret i32 -1\n}\n")
|
||||
|
||||
# chars(s) -> []int : a fresh %LSlice of every code point, in order
|
||||
emith("define ptr @fn_uni_chars(ptr %s) {\n")
|
||||
emith("entry:\n %n = call i32 @fn_uni_len(ptr %s)\n %h = call ptr @malloc(i64 16)\n")
|
||||
emith(" %nz = zext i32 %n to i64\n %bytes = mul i64 %nz, 4\n %data = call ptr @malloc(i64 %bytes)\n")
|
||||
emith(" %d0 = getelementptr inbounds %LSlice, ptr %h, i32 0, i32 0\n store ptr %data, ptr %d0\n")
|
||||
emith(" %d1 = getelementptr inbounds %LSlice, ptr %h, i32 0, i32 1\n store i32 %n, ptr %d1\n")
|
||||
emith(" %d2 = getelementptr inbounds %LSlice, ptr %h, i32 0, i32 2\n store i32 %n, ptr %d2\n")
|
||||
emith(" %ip = alloca i32\n %kp = alloca i32\n %cp = alloca i32\n store i32 0, ptr %ip\n store i32 0, ptr %kp\n br label %lp\n")
|
||||
emith("lp:\n %i = load i32, ptr %ip\n %p = getelementptr i8, ptr %s, i32 %i\n %b = load i8, ptr %p\n %c = zext i8 %b to i32\n")
|
||||
emith(" %z = icmp eq i32 %c, 0\n br i1 %z, label %done, label %go\n")
|
||||
emith("go:\n %ni = call i32 @fn_uni_decode(ptr %s, i32 %i, ptr %cp)\n %v = load i32, ptr %cp\n %k = load i32, ptr %kp\n")
|
||||
emith(" %slot = getelementptr i32, ptr %data, i32 %k\n store i32 %v, ptr %slot\n %k1 = add i32 %k, 1\n store i32 %k1, ptr %kp\n store i32 %ni, ptr %ip\n br label %lp\n")
|
||||
emith("done:\n ret ptr %h\n}\n")
|
||||
|
||||
# case map the whole string. %up != 0 -> uppercase, else lowercase. Decodes,
|
||||
# maps each code point, and re-encodes into a fresh buffer (worst case 4 bytes
|
||||
# per code point, though ASCII/Latin-1 mapping preserves byte length).
|
||||
emith("define ptr @fn_uni_case(ptr %s, i32 %up) {\n")
|
||||
emith("entry:\n %bl = call i64 @strlen(ptr %s)\n %cap0 = mul i64 %bl, 4\n %cap = add i64 %cap0, 4\n %out = call ptr @malloc(i64 %cap)\n")
|
||||
emith(" %ip = alloca i32\n %op = alloca i32\n %cp = alloca i32\n store i32 0, ptr %ip\n store i32 0, ptr %op\n br label %lp\n")
|
||||
emith("lp:\n %i = load i32, ptr %ip\n %p = getelementptr i8, ptr %s, i32 %i\n %b = load i8, ptr %p\n %c = zext i8 %b to i32\n")
|
||||
emith(" %z = icmp eq i32 %c, 0\n br i1 %z, label %done, label %go\n")
|
||||
emith("go:\n %ni = call i32 @fn_uni_decode(ptr %s, i32 %i, ptr %cp)\n %v = load i32, ptr %cp\n")
|
||||
emith(" %mu = call i32 @fn_uni_upcp(i32 %v)\n %ml = call i32 @fn_uni_locp(i32 %v)\n %isup = icmp ne i32 %up, 0\n %m = select i1 %isup, i32 %mu, i32 %ml\n")
|
||||
emith(" %o = load i32, ptr %op\n %no = call i32 @fn_uni_encode(ptr %out, i32 %o, i32 %m)\n store i32 %no, ptr %op\n store i32 %ni, ptr %ip\n br label %lp\n")
|
||||
emith("done:\n %fo = load i32, ptr %op\n %endp = getelementptr i8, ptr %out, i32 %fo\n store i8 0, ptr %endp\n ret ptr %out\n}\n")
|
||||
|
||||
# truncate(s, n) -> the first n code points as a fresh string (never splits a
|
||||
# multibyte character). n <= 0 yields the empty string.
|
||||
emith("define ptr @fn_uni_truncate(ptr %s, i32 %n) {\n")
|
||||
emith("entry:\n %ip = alloca i32\n %kp = alloca i32\n %cp = alloca i32\n store i32 0, ptr %ip\n store i32 0, ptr %kp\n br label %lp\n")
|
||||
emith("lp:\n %k = load i32, ptr %kp\n %enough = icmp sge i32 %k, %n\n br i1 %enough, label %cut, label %chk\n")
|
||||
emith("chk:\n %i = load i32, ptr %ip\n %p = getelementptr i8, ptr %s, i32 %i\n %b = load i8, ptr %p\n %c = zext i8 %b to i32\n")
|
||||
emith(" %z = icmp eq i32 %c, 0\n br i1 %z, label %cut, label %go\n")
|
||||
emith("go:\n %ni = call i32 @fn_uni_decode(ptr %s, i32 %i, ptr %cp)\n %k1 = add i32 %k, 1\n store i32 %k1, ptr %kp\n store i32 %ni, ptr %ip\n br label %lp\n")
|
||||
emith("cut:\n %len = load i32, ptr %ip\n %lz = zext i32 %len to i64\n %cap = add i64 %lz, 1\n %out = call ptr @malloc(i64 %cap)\n %lz2 = zext i32 %len to i64\n call ptr @memcpy(ptr %out, ptr %s, i64 %lz2)\n")
|
||||
emith(" %endp = getelementptr i8, ptr %out, i32 %len\n store i8 0, ptr %endp\n ret ptr %out\n}\n")
|
||||
|
||||
emit_uni_grapheme()
|
||||
}
|
||||
|
||||
# grapheme_len: an approximate UAX#29 count. A new cluster starts on each code
|
||||
# point except: a combining/variation "extend"; the code point after a ZWJ (so
|
||||
# ZWJ emoji sequences count as one); and the second regional indicator of a flag
|
||||
# pair. Simple state carried in allocas.
|
||||
function emit_uni_grapheme() -> void {
|
||||
emith("define i32 @fn_uni_grapheme_len(ptr %s) {\n")
|
||||
emith("entry:\n %ip = alloca i32\n %np = alloca i32\n %zp = alloca i32\n %rp = alloca i32\n %cp = alloca i32\n")
|
||||
emith(" store i32 0, ptr %ip\n store i32 0, ptr %np\n store i32 0, ptr %zp\n store i32 0, ptr %rp\n br label %lp\n")
|
||||
emith("lp:\n %i = load i32, ptr %ip\n %p = getelementptr i8, ptr %s, i32 %i\n %b = load i8, ptr %p\n %c = zext i8 %b to i32\n")
|
||||
emith(" %z = icmp eq i32 %c, 0\n br i1 %z, label %done, label %go\n")
|
||||
emith("go:\n %ni = call i32 @fn_uni_decode(ptr %s, i32 %i, ptr %cp)\n %v = load i32, ptr %cp\n store i32 %ni, ptr %ip\n")
|
||||
# ZWJ (U+200D): extends the cluster and arms the join for the next code point
|
||||
emith(" %iszwj = icmp eq i32 %v, 8205\n br i1 %iszwj, label %zwj, label %notzwj\n")
|
||||
emith("zwj:\n %n0 = load i32, ptr %np\n %n0z = icmp eq i32 %n0, 0\n %n0b = zext i1 %n0z to i32\n %n0n = add i32 %n0, %n0b\n store i32 %n0n, ptr %np\n") # leading ZWJ still opens one cluster
|
||||
emith(" store i32 1, ptr %zp\n store i32 0, ptr %rp\n br label %lp\n")
|
||||
emith("notzwj:\n %n = load i32, ptr %np\n %first = icmp eq i32 %n, 0\n br i1 %first, label %open, label %cont\n")
|
||||
# first cluster
|
||||
emith("open:\n store i32 1, ptr %np\n store i32 0, ptr %zp\n %ri0 = call i32 @fn_uni_inrange(i32 %v, i32 127462, i32 127487)\n store i32 %ri0, ptr %rp\n br label %lp\n")
|
||||
emith("cont:\n %zj = load i32, ptr %zp\n %afterz = icmp ne i32 %zj, 0\n br i1 %afterz, label %joinz, label %chkext\n")
|
||||
# code point right after a ZWJ joins the current cluster
|
||||
emith("joinz:\n store i32 0, ptr %zp\n store i32 0, ptr %rp\n br label %lp\n")
|
||||
emith("chkext:\n %ext = call i32 @fn_uni_is_extend(i32 %v)\n %isext = icmp ne i32 %ext, 0\n br i1 %isext, label %joinext, label %chkri\n")
|
||||
emith("joinext:\n store i32 0, ptr %rp\n br label %lp\n")
|
||||
# regional indicator: joins only as the second of a pair
|
||||
emith("chkri:\n %ri = call i32 @fn_uni_inrange(i32 %v, i32 127462, i32 127487)\n %isri = icmp ne i32 %ri, 0\n %ropen = load i32, ptr %rp\n %ropenb = icmp ne i32 %ropen, 0\n %pair = and i1 %isri, %ropenb\n br i1 %pair, label %joinri, label %newcl\n")
|
||||
emith("joinri:\n store i32 0, ptr %rp\n br label %lp\n")
|
||||
emith("newcl:\n %nn = load i32, ptr %np\n %nn1 = add i32 %nn, 1\n store i32 %nn1, ptr %np\n store i32 0, ptr %zp\n %riset = select i1 %isri, i32 1, i32 0\n store i32 %riset, ptr %rp\n br label %lp\n")
|
||||
emith("done:\n %r = load i32, ptr %np\n ret i32 %r\n}\n")
|
||||
}
|
||||
21774
selfhost/ludicc.seed.ll
21774
selfhost/ludicc.seed.ll
File diff suppressed because it is too large
Load diff
|
|
@ -457,5 +457,16 @@
|
|||
"os-config_dir",
|
||||
"os-cache_dir",
|
||||
"os-temp_dir"
|
||||
],
|
||||
"unicode": [
|
||||
"unicode-len",
|
||||
"unicode-byte_len",
|
||||
"unicode-is_valid_utf8",
|
||||
"unicode-char_at",
|
||||
"unicode-chars",
|
||||
"unicode-upper",
|
||||
"unicode-lower",
|
||||
"unicode-truncate",
|
||||
"unicode-grapheme_len"
|
||||
]
|
||||
}
|
||||
|
|
|
|||
|
|
@ -33,6 +33,7 @@ function selfhost_frags() -> []pointer {
|
|||
push(f, "selfhost/emit_noise.ludic")
|
||||
push(f, "selfhost/emit_log.ludic")
|
||||
push(f, "selfhost/emit_os.ludic")
|
||||
push(f, "selfhost/emit_unicode.ludic")
|
||||
push(f, "selfhost/emit_list.ludic")
|
||||
push(f, "selfhost/emit_ease.ludic")
|
||||
push(f, "selfhost/emit_collide.ludic")
|
||||
|
|
|
|||
|
|
@ -104,6 +104,7 @@ function cmd_test() -> int {
|
|||
feat_case("library/noise", "", "1 2 3 4 5 6 7 8 9 10 11", "noise.ludic (Noise value/perlin/simplex/fbm/cellular determinism + range)")
|
||||
feat_case("library/logging", "", "0 5 2 1", "logging.ludic (Log levels, set_level/level threshold, structured fields)")
|
||||
feat_case("library/os", "", "1 2 3 4 5 6 7 8 9 10 11 12 13", "os.ludic (Os args/env round-trip, platform/arch, known dirs)")
|
||||
feat_case("library/unicode", "", "1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18", "unicode.ludic (UTF-8 len/validate/char_at/chars/case/truncate/graphemes)")
|
||||
|
||||
# issue #9: the Time/Date/Duration/Clock stdlib, driven from its own `entry`.
|
||||
net_case("lang/offline_rewards", "13 650 2026-08-30 0")
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue