feat(stdlib): add Regex.* — a linear-time regular-expression engine (#18)
A regular-expression library with PCRE/PECL-compatible syntax, implemented as a
Thompson NFA / Pike VM so a bad pattern from a modder can NEVER cause
catastrophic backtracking — matching is O(n·m), never exponential. `(a+)+$` on
40 non-matching chars, `(a*)*b`, `(.*a){20}b` all run in microseconds; a 50 KB
input scans in ~7 ms.
The engine (runtime/native/regex.ludic + regex_vm.ludic, ~700 lines of Ludic, no
C) parses a pattern to a small bytecode program — an unanchored lazy `.*?` prefix
makes a plain search match anywhere — and the VM runs every alive thread in
lockstep per input byte, deduped by program counter and carrying capture slots
(save/restore, leftmost-greedy priority). Supported: literals, `.`, classes
`[...]` (ranges, negation, `\d \w \s` and their negations), anchors `^ $`,
alternation `|`, capturing and `(?:…)` groups, and `* + ? {n} {n,} {n,m}` in
greedy or lazy form, plus the common escapes; numbered capture groups. Errors are
values — an invalid pattern compiles to null, never a crash. Backreferences and
look-around are out of scope for a linear engine, and on the degenerate case of a
nullable subpattern under an unbounded quantifier positions may differ from a
backtracking engine (the price of the linear-time guarantee) — documented.
Surface (Regex.*, aliased in emit_call.ludic to the regex_* functions):
compile / valid / matches / test / find / exec / next / replace / group /
group_count / start / end / ok.
The runtime is spliced on demand: the parser sets a flag when it sees `Regex.`
and maybe_splice_runtime imports the engine — so it costs nothing in a program
that doesn't use it and works in a plain tool (not just an ECS game).
Verified against Python's `re` as an oracle: a 20k-case grammar fuzzer agrees
100% on realistic patterns (0 / 15000 with capture groups) and 99.8% on group-0
spans across the full pathological grammar, the residual being the documented
nullable-quantifier case. examples/library/regex.ludic asserts the behaviour
(wired into `x test`, now 60 passed); docs: a Regex section + 13 per-symbol
pages, inventory + coverage green. Seed reseeded; the C-free bootstrap fixpoint
holds.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
9ed0070039
commit
b798e3024e
22 changed files with 15079 additions and 13041 deletions
27
docs/language/regex/regex-end.md
Normal file
27
docs/language/regex/regex-end.md
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
---
|
||||
id: regex-end
|
||||
name: Regex.end
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.end
|
||||
sig: Regex.end(match, n) -> int
|
||||
tip: The end byte offset of group n, or -1 if it did not participate.
|
||||
order: 12
|
||||
ns: Regex
|
||||
member: end
|
||||
---
|
||||
|
||||
Returns the byte offset just past the end of group <code>n</code> of <code>match</code>, or <code>-1</code> if the group did not participate. Feed <code>Regex.end(m, 0)</code> back into <a href="regex-next"><code>Regex.next</code></a> to walk every match.
|
||||
|
||||
Parameters:
|
||||
- `match` — a `Match` value
|
||||
- `n` — the group number
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
let m = Regex.find("hi there", "hi")
|
||||
print(Regex.end(m, 0))
|
||||
}
|
||||
}
|
||||
```
|
||||
Loading…
Add table
Add a link
Reference in a new issue