ludic/examples/library/regex.ludic
Orkuncakilkaya b798e3024e
All checks were successful
bootstrap / cfree-fixpoint (push) Successful in 26s
ci / build-and-test (push) Successful in 1m6s
commit-lint / conventional-commits (push) Successful in 4s
docs / build-and-deploy (push) Successful in 18s
feat(stdlib): add Regex.* — a linear-time regular-expression engine (#18)
A regular-expression library with PCRE/PECL-compatible syntax, implemented as a
Thompson NFA / Pike VM so a bad pattern from a modder can NEVER cause
catastrophic backtracking — matching is O(n·m), never exponential. `(a+)+$` on
40 non-matching chars, `(a*)*b`, `(.*a){20}b` all run in microseconds; a 50 KB
input scans in ~7 ms.

The engine (runtime/native/regex.ludic + regex_vm.ludic, ~700 lines of Ludic, no
C) parses a pattern to a small bytecode program — an unanchored lazy `.*?` prefix
makes a plain search match anywhere — and the VM runs every alive thread in
lockstep per input byte, deduped by program counter and carrying capture slots
(save/restore, leftmost-greedy priority). Supported: literals, `.`, classes
`[...]` (ranges, negation, `\d \w \s` and their negations), anchors `^ $`,
alternation `|`, capturing and `(?:…)` groups, and `* + ? {n} {n,} {n,m}` in
greedy or lazy form, plus the common escapes; numbered capture groups. Errors are
values — an invalid pattern compiles to null, never a crash. Backreferences and
look-around are out of scope for a linear engine, and on the degenerate case of a
nullable subpattern under an unbounded quantifier positions may differ from a
backtracking engine (the price of the linear-time guarantee) — documented.

Surface (Regex.*, aliased in emit_call.ludic to the regex_* functions):
compile / valid / matches / test / find / exec / next / replace / group /
group_count / start / end / ok.

The runtime is spliced on demand: the parser sets a flag when it sees `Regex.`
and maybe_splice_runtime imports the engine — so it costs nothing in a program
that doesn't use it and works in a plain tool (not just an ECS game).

Verified against Python's `re` as an oracle: a 20k-case grammar fuzzer agrees
100% on realistic patterns (0 / 15000 with capture groups) and 99.8% on group-0
spans across the full pathological grammar, the residual being the documented
nullable-quantifier case. examples/library/regex.ludic asserts the behaviour
(wired into `x test`, now 60 passed); docs: a Regex section + 13 per-symbol
pages, inventory + coverage green. Seed reseeded; the C-free bootstrap fixpoint
holds.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-31 12:23:19 +03:00

59 lines
2.5 KiB
Text

# regex.ludic — Regex.* behaviour and the linear-time safety guarantee. Each
# assertion that holds prints its number, so a full run prints:
# 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18
# The engine is a Thompson NFA / Pike VM (see runtime/native/regex.ludic): a
# pathological pattern is O(n·m), never exponential — check #17 pins that.
program Regex {
entry {
# --- search (matches) ---
if Regex.matches("hello world", "wor.d") { print(1) }
if not Regex.matches("hello", "\\d+") { print(2) }
# --- anchors + alternation + groups (a chat-command validator) ---
if Regex.matches("/help", "^/(help|quit)$") { print(3) }
if not Regex.matches("/helpme", "^/(help|quit)$") { print(4) }
# --- character classes / shorthands ---
if Regex.matches("abc123", "\\d+") { print(5) }
if Regex.matches("user_1@host.io", "^\\w+@\\w+\\.\\w+$") { print(6) }
if Regex.matches("x", "[a-z]") { print(7) }
if not Regex.matches("5", "[^0-9]") { print(8) }
# --- greedy vs lazy ---
let g = Regex.find("<a><b>", "<(.+)>")
if Regex.group(g, 1) == "a><b" { print(9) }
let l = Regex.find("<a><b>", "<(.+?)>")
if Regex.group(l, 1) == "a" { print(10) }
# --- bounded repetition {n,m} ---
if Regex.matches("aaa", "^a{2,3}$") { print(11) }
if not Regex.matches("aaaa", "^a{2,3}$") { print(12) }
# --- capture groups by number ---
let m = Regex.find("name: Alice", "(\\w+): (\\w+)")
if Regex.group(m, 1) == "name" and Regex.group(m, 2) == "Alice" { print(13) }
# --- replace: literal and with a \1 backreference (dialogue tokens) ---
if Regex.replace("a1b2c3", "\\d", "#") == "a#b#c#" { print(14) }
if Regex.replace("hi {name}!", "\\{(\\w+)\\}", "[\\1]") == "hi [name]!" { print(15) }
# --- find_all via a compile-once / next loop (extract #hashtags) ---
let re = Regex.compile("#(\\w+)")
var tags = ""
var cur = Regex.exec("#a #bb #ccc", re)
while Regex.ok(cur) {
tags = tags + Regex.group(cur, 1)
cur = Regex.next("#a #bb #ccc", re, Regex.end(cur, 0))
}
if tags == "abbccc" { print(16) }
# --- linear time: a classic catastrophic-backtracking pattern still runs
# (a naive engine would hang on 40 a's; the NFA is O(n·m)) ---
if Regex.matches("aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa", "(a+)+$") { print(17) }
# --- errors are values: an invalid pattern compiles to null, never crashes ---
if not Regex.valid("(unterminated") { print(18) }
exit(0)
}
}