feat(stdlib): add Regex.* — a linear-time regular-expression engine (#18)
All checks were successful
bootstrap / cfree-fixpoint (push) Successful in 26s
ci / build-and-test (push) Successful in 1m6s
commit-lint / conventional-commits (push) Successful in 4s
docs / build-and-deploy (push) Successful in 18s

A regular-expression library with PCRE/PECL-compatible syntax, implemented as a
Thompson NFA / Pike VM so a bad pattern from a modder can NEVER cause
catastrophic backtracking — matching is O(n·m), never exponential. `(a+)+$` on
40 non-matching chars, `(a*)*b`, `(.*a){20}b` all run in microseconds; a 50 KB
input scans in ~7 ms.

The engine (runtime/native/regex.ludic + regex_vm.ludic, ~700 lines of Ludic, no
C) parses a pattern to a small bytecode program — an unanchored lazy `.*?` prefix
makes a plain search match anywhere — and the VM runs every alive thread in
lockstep per input byte, deduped by program counter and carrying capture slots
(save/restore, leftmost-greedy priority). Supported: literals, `.`, classes
`[...]` (ranges, negation, `\d \w \s` and their negations), anchors `^ $`,
alternation `|`, capturing and `(?:…)` groups, and `* + ? {n} {n,} {n,m}` in
greedy or lazy form, plus the common escapes; numbered capture groups. Errors are
values — an invalid pattern compiles to null, never a crash. Backreferences and
look-around are out of scope for a linear engine, and on the degenerate case of a
nullable subpattern under an unbounded quantifier positions may differ from a
backtracking engine (the price of the linear-time guarantee) — documented.

Surface (Regex.*, aliased in emit_call.ludic to the regex_* functions):
compile / valid / matches / test / find / exec / next / replace / group /
group_count / start / end / ok.

The runtime is spliced on demand: the parser sets a flag when it sees `Regex.`
and maybe_splice_runtime imports the engine — so it costs nothing in a program
that doesn't use it and works in a plain tool (not just an ECS game).

Verified against Python's `re` as an oracle: a 20k-case grammar fuzzer agrees
100% on realistic patterns (0 / 15000 with capture groups) and 99.8% on group-0
spans across the full pathological grammar, the residual being the documented
nullable-quantifier case. examples/library/regex.ludic asserts the behaviour
(wired into `x test`, now 60 passed); docs: a Regex section + 13 per-symbol
pages, inventory + coverage green. Seed reseeded; the C-free bootstrap fixpoint
holds.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Orkun ÇAKILKAYA 2026-08-31 12:23:19 +03:00
parent 9ed0070039
commit b798e3024e
22 changed files with 15079 additions and 13041 deletions

View file

@ -0,0 +1,7 @@
---
id: regex
title: Regex
order: 26
---
Regular expressions with PCRE/PECL-compatible syntax over a **linear-time** Thompson NFA (a Pike VM), so a bad pattern from a modder can never trigger catastrophic backtracking — matching is <code>O(n·m)</code>, never exponential. Compile a pattern once and reuse it; an invalid pattern is an <a href="regex-compile"><code>error value</code></a>, never a crash. Supported: literals, <code>.</code>, classes <code>[...]</code> with ranges/negation and <code>\d \w \s</code> (and their negations), anchors <code>^ $</code>, alternation <code>|</code>, capturing and <code>(?:…)</code> groups, and the quantifiers <code>* + ?</code> and <code>{n,m}</code> in greedy or lazy form. Backreferences and look-around are out of scope for a linear engine.

View file

@ -0,0 +1,26 @@
---
id: regex-compile
name: Regex.compile
category: regex
kind: namespace-method
tokens: Regex.compile
sig: Regex.compile(pattern) -> Regex
tip: Compile a pattern once for reuse; returns an error value (null) if it is invalid.
order: 1
ns: Regex
member: compile
---
Compiles <code>pattern</code> to a reusable <code>Regex</code> and returns it, or <code>null</code> if the pattern is malformed — so a bad pattern is a value you can check, never a crash. Compiling once and reusing with <a href="regex-test"><code>Regex.test</code></a> / <a href="regex-exec"><code>Regex.exec</code></a> avoids re-parsing the pattern every call, which matters in a hot loop.
Parameters:
- `pattern` — the regular expression source
```ludic
program Demo {
entry {
let re = Regex.compile("^\\d{4}-\\d{2}-\\d{2}$")
if Regex.test("2026-08-31", re) { print(1) }
}
}
```

View file

@ -0,0 +1,27 @@
---
id: regex-end
name: Regex.end
category: regex
kind: namespace-method
tokens: Regex.end
sig: Regex.end(match, n) -> int
tip: The end byte offset of group n, or -1 if it did not participate.
order: 12
ns: Regex
member: end
---
Returns the byte offset just past the end of group <code>n</code> of <code>match</code>, or <code>-1</code> if the group did not participate. Feed <code>Regex.end(m, 0)</code> back into <a href="regex-next"><code>Regex.next</code></a> to walk every match.
Parameters:
- `match` — a `Match` value
- `n` — the group number
```ludic
program Demo {
entry {
let m = Regex.find("hi there", "hi")
print(Regex.end(m, 0))
}
}
```

View file

@ -0,0 +1,28 @@
---
id: regex-exec
name: Regex.exec
category: regex
kind: namespace-method
tokens: Regex.exec
sig: Regex.exec(text, re) -> Match
tip: Like find, but with an already-compiled Regex.
order: 6
ns: Regex
member: exec
---
The compiled-pattern form of <a href="regex-find"><code>Regex.find</code></a>: returns the first <code>Match</code> of the pre-compiled <code>re</code> in <code>text</code>, or <code>null</code>.
Parameters:
- `text` — the string to search
- `re` — a compiled `Regex`
```ludic
program Demo {
entry {
let re = Regex.compile("#(\\w+)")
let m = Regex.exec("a #tag b", re)
print(Regex.group(m, 1))
}
}
```

View file

@ -0,0 +1,27 @@
---
id: regex-find
name: Regex.find
category: regex
kind: namespace-method
tokens: Regex.find
sig: Regex.find(text, pattern) -> Match
tip: The first match of the pattern in the text (a Match, or null if none).
order: 5
ns: Regex
member: find
---
Searches <code>text</code> and returns the first <code>Match</code> — carrying the overall span and every capture group — or <code>null</code> if the pattern does not match. Read groups with <a href="regex-group"><code>Regex.group</code></a> and spans with <a href="regex-start"><code>Regex.start</code></a> / <a href="regex-end"><code>Regex.end</code></a>.
Parameters:
- `text` — the string to search
- `pattern` — the regular expression source
```ludic
program Demo {
entry {
let m = Regex.find("name: Alice", "(\\w+): (\\w+)")
if Regex.ok(m) { print(1) }
}
}
```

View file

@ -0,0 +1,27 @@
---
id: regex-group
name: Regex.group
category: regex
kind: namespace-method
tokens: Regex.group
sig: Regex.group(match, n) -> string
tip: The text of capture group n (group 0 is the whole match); empty if unset.
order: 9
ns: Regex
member: group
---
Returns the substring captured by group <code>n</code> of <code>match</code> — group <code>0</code> is the whole match, <code>1</code> the first parenthesized group, and so on. Returns an empty string if the group did not participate in the match.
Parameters:
- `match` — a `Match` from `Regex.find` / `Regex.exec` / `Regex.next`
- `n` — the group number (0 = whole match)
```ludic
program Demo {
entry {
let m = Regex.find("key=value", "(\\w+)=(\\w+)")
print(Regex.group(m, 2))
}
}
```

View file

@ -0,0 +1,26 @@
---
id: regex-group_count
name: Regex.group_count
category: regex
kind: namespace-method
tokens: Regex.group_count
sig: Regex.group_count(match) -> int
tip: How many capture groups the match's pattern has.
order: 10
ns: Regex
member: group_count
---
Returns the number of capture groups in the pattern behind <code>match</code> (not counting group 0, the whole match). Valid group numbers for <a href="regex-group"><code>Regex.group</code></a> are <code>0</code>…<code>group_count</code>.
Parameters:
- `match` — a `Match` value
```ludic
program Demo {
entry {
let m = Regex.find("a-b", "(\\w)-(\\w)")
print(Regex.group_count(m))
}
}
```

View file

@ -0,0 +1,26 @@
---
id: regex-matches
name: Regex.matches
category: regex
kind: namespace-method
tokens: Regex.matches
sig: Regex.matches(text, pattern) -> bool
tip: True if the pattern matches anywhere in the text (a search).
order: 3
ns: Regex
member: matches
---
Searches <code>text</code> for <code>pattern</code> and returns whether it matches <em>anywhere</em> (like a search, not a full-string match — anchor with <code>^…$</code> for that). Compiles the pattern each call; for a pattern you use repeatedly, compile once with <a href="regex-compile"><code>Regex.compile</code></a> and use <a href="regex-test"><code>Regex.test</code></a>.
Parameters:
- `text` — the string to search
- `pattern` — the regular expression source
```ludic
program Demo {
entry {
if Regex.matches("/quit now", "^/(help|quit)\\b") { print(1) }
}
}
```

View file

@ -0,0 +1,32 @@
---
id: regex-next
name: Regex.next
category: regex
kind: namespace-method
tokens: Regex.next
sig: Regex.next(text, re, from) -> Match
tip: The next match at or after byte offset from — the basis of a find-all loop.
order: 7
ns: Regex
member: next
---
Returns the first match of <code>re</code> in <code>text</code> starting at or after byte offset <code>from</code>, or <code>null</code>. Loop it from the previous match's end (see <a href="regex-end"><code>Regex.end</code></a>) to iterate every match — the find-all idiom.
Parameters:
- `text` — the string to search
- `re` — a compiled `Regex`
- `from` — the byte offset to resume from
```ludic
program Demo {
entry {
let re = Regex.compile("#(\\w+)")
var m = Regex.exec("#a #b #c", re)
while Regex.ok(m) {
print(Regex.group(m, 1))
m = Regex.next("#a #b #c", re, Regex.end(m, 0))
}
}
}
```

View file

@ -0,0 +1,25 @@
---
id: regex-ok
name: Regex.ok
category: regex
kind: namespace-method
tokens: Regex.ok
sig: Regex.ok(match) -> bool
tip: True if the match succeeded (the Match is non-null).
order: 13
ns: Regex
member: ok
---
Returns whether <code>match</code> is a real match rather than <code>null</code> — the readable way to test the result of <a href="regex-find"><code>Regex.find</code></a> / <a href="regex-exec"><code>Regex.exec</code></a> / <a href="regex-next"><code>Regex.next</code></a>, and the loop condition for a find-all.
Parameters:
- `match` — a `Match` value (possibly null)
```ludic
program Demo {
entry {
if not Regex.ok(Regex.find("abc", "\\d")) { print(1) }
}
}
```

View file

@ -0,0 +1,27 @@
---
id: regex-replace
name: Regex.replace
category: regex
kind: namespace-method
tokens: Regex.replace
sig: Regex.replace(text, pattern, replacement) -> string
tip: Replace every match; the replacement expands \0..\9 group references.
order: 8
ns: Regex
member: replace
---
Returns a new string with every non-overlapping match of <code>pattern</code> in <code>text</code> replaced by <code>replacement</code>. In the replacement, <code>\0</code> is the whole match, <code>\1</code>…<code>\9</code> are capture groups, and <code>\\</code> is a literal backslash — the natural fit for dialogue/localization token substitution.
Parameters:
- `text` — the string to transform
- `pattern` — the regular expression source
- `replacement` — the replacement, with `\0`..`\9` group references
```ludic
program Demo {
entry {
print(Regex.replace("hi {name}!", "\\{(\\w+)\\}", "[\\1]"))
}
}
```

View file

@ -0,0 +1,27 @@
---
id: regex-start
name: Regex.start
category: regex
kind: namespace-method
tokens: Regex.start
sig: Regex.start(match, n) -> int
tip: The start byte offset of group n, or -1 if it did not participate.
order: 11
ns: Regex
member: start
---
Returns the byte offset where group <code>n</code> of <code>match</code> begins, or <code>-1</code> if the group did not participate. Group <code>0</code> is the whole match, so <code>Regex.start(m, 0)</code> is where the match begins.
Parameters:
- `match` — a `Match` value
- `n` — the group number
```ludic
program Demo {
entry {
let m = Regex.find(" hi", "hi")
print(Regex.start(m, 0))
}
}
```

View file

@ -0,0 +1,27 @@
---
id: regex-test
name: Regex.test
category: regex
kind: namespace-method
tokens: Regex.test
sig: Regex.test(text, re) -> bool
tip: Like matches, but reuses an already-compiled Regex.
order: 4
ns: Regex
member: test
---
The compiled-pattern form of <a href="regex-matches"><code>Regex.matches</code></a>: returns whether the pre-compiled <code>re</code> matches anywhere in <code>text</code>. No per-call parsing, so this is the one to call every frame.
Parameters:
- `text` — the string to search
- `re` — a compiled `Regex` from `Regex.compile`
```ludic
program Demo {
entry {
let re = Regex.compile("[a-z]+")
if Regex.test("hello", re) { print(1) }
}
}
```

View file

@ -0,0 +1,26 @@
---
id: regex-valid
name: Regex.valid
category: regex
kind: namespace-method
tokens: Regex.valid
sig: Regex.valid(pattern) -> bool
tip: True if the pattern is well-formed (compiles without error).
order: 2
ns: Regex
member: valid
---
A convenience over <a href="regex-compile"><code>Regex.compile</code></a>: <code>true</code> when <code>pattern</code> compiles, <code>false</code> when it is malformed. Use it to validate a user- or mod-supplied pattern before relying on it.
Parameters:
- `pattern` — the regular expression source
```ludic
program Demo {
entry {
if Regex.valid("(a|b)+") { print(1) }
if not Regex.valid("(unterminated") { print(2) }
}
}
```