feat(stdlib): add Regex.* — a linear-time regular-expression engine (#18)
A regular-expression library with PCRE/PECL-compatible syntax, implemented as a
Thompson NFA / Pike VM so a bad pattern from a modder can NEVER cause
catastrophic backtracking — matching is O(n·m), never exponential. `(a+)+$` on
40 non-matching chars, `(a*)*b`, `(.*a){20}b` all run in microseconds; a 50 KB
input scans in ~7 ms.
The engine (runtime/native/regex.ludic + regex_vm.ludic, ~700 lines of Ludic, no
C) parses a pattern to a small bytecode program — an unanchored lazy `.*?` prefix
makes a plain search match anywhere — and the VM runs every alive thread in
lockstep per input byte, deduped by program counter and carrying capture slots
(save/restore, leftmost-greedy priority). Supported: literals, `.`, classes
`[...]` (ranges, negation, `\d \w \s` and their negations), anchors `^ $`,
alternation `|`, capturing and `(?:…)` groups, and `* + ? {n} {n,} {n,m}` in
greedy or lazy form, plus the common escapes; numbered capture groups. Errors are
values — an invalid pattern compiles to null, never a crash. Backreferences and
look-around are out of scope for a linear engine, and on the degenerate case of a
nullable subpattern under an unbounded quantifier positions may differ from a
backtracking engine (the price of the linear-time guarantee) — documented.
Surface (Regex.*, aliased in emit_call.ludic to the regex_* functions):
compile / valid / matches / test / find / exec / next / replace / group /
group_count / start / end / ok.
The runtime is spliced on demand: the parser sets a flag when it sees `Regex.`
and maybe_splice_runtime imports the engine — so it costs nothing in a program
that doesn't use it and works in a plain tool (not just an ECS game).
Verified against Python's `re` as an oracle: a 20k-case grammar fuzzer agrees
100% on realistic patterns (0 / 15000 with capture groups) and 99.8% on group-0
spans across the full pathological grammar, the residual being the documented
nullable-quantifier case. examples/library/regex.ludic asserts the behaviour
(wired into `x test`, now 60 passed); docs: a Regex section + 13 per-symbol
pages, inventory + coverage green. Seed reseeded; the C-free bootstrap fixpoint
holds.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
9ed0070039
commit
b798e3024e
22 changed files with 15079 additions and 13041 deletions
7
docs/language/regex/_section.md
Normal file
7
docs/language/regex/_section.md
Normal file
|
|
@ -0,0 +1,7 @@
|
|||
---
|
||||
id: regex
|
||||
title: Regex
|
||||
order: 26
|
||||
---
|
||||
|
||||
Regular expressions with PCRE/PECL-compatible syntax over a **linear-time** Thompson NFA (a Pike VM), so a bad pattern from a modder can never trigger catastrophic backtracking — matching is <code>O(n·m)</code>, never exponential. Compile a pattern once and reuse it; an invalid pattern is an <a href="regex-compile"><code>error value</code></a>, never a crash. Supported: literals, <code>.</code>, classes <code>[...]</code> with ranges/negation and <code>\d \w \s</code> (and their negations), anchors <code>^ $</code>, alternation <code>|</code>, capturing and <code>(?:…)</code> groups, and the quantifiers <code>* + ?</code> and <code>{n,m}</code> in greedy or lazy form. Backreferences and look-around are out of scope for a linear engine.
|
||||
26
docs/language/regex/regex-compile.md
Normal file
26
docs/language/regex/regex-compile.md
Normal file
|
|
@ -0,0 +1,26 @@
|
|||
---
|
||||
id: regex-compile
|
||||
name: Regex.compile
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.compile
|
||||
sig: Regex.compile(pattern) -> Regex
|
||||
tip: Compile a pattern once for reuse; returns an error value (null) if it is invalid.
|
||||
order: 1
|
||||
ns: Regex
|
||||
member: compile
|
||||
---
|
||||
|
||||
Compiles <code>pattern</code> to a reusable <code>Regex</code> and returns it, or <code>null</code> if the pattern is malformed — so a bad pattern is a value you can check, never a crash. Compiling once and reusing with <a href="regex-test"><code>Regex.test</code></a> / <a href="regex-exec"><code>Regex.exec</code></a> avoids re-parsing the pattern every call, which matters in a hot loop.
|
||||
|
||||
Parameters:
|
||||
- `pattern` — the regular expression source
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
let re = Regex.compile("^\\d{4}-\\d{2}-\\d{2}$")
|
||||
if Regex.test("2026-08-31", re) { print(1) }
|
||||
}
|
||||
}
|
||||
```
|
||||
27
docs/language/regex/regex-end.md
Normal file
27
docs/language/regex/regex-end.md
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
---
|
||||
id: regex-end
|
||||
name: Regex.end
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.end
|
||||
sig: Regex.end(match, n) -> int
|
||||
tip: The end byte offset of group n, or -1 if it did not participate.
|
||||
order: 12
|
||||
ns: Regex
|
||||
member: end
|
||||
---
|
||||
|
||||
Returns the byte offset just past the end of group <code>n</code> of <code>match</code>, or <code>-1</code> if the group did not participate. Feed <code>Regex.end(m, 0)</code> back into <a href="regex-next"><code>Regex.next</code></a> to walk every match.
|
||||
|
||||
Parameters:
|
||||
- `match` — a `Match` value
|
||||
- `n` — the group number
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
let m = Regex.find("hi there", "hi")
|
||||
print(Regex.end(m, 0))
|
||||
}
|
||||
}
|
||||
```
|
||||
28
docs/language/regex/regex-exec.md
Normal file
28
docs/language/regex/regex-exec.md
Normal file
|
|
@ -0,0 +1,28 @@
|
|||
---
|
||||
id: regex-exec
|
||||
name: Regex.exec
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.exec
|
||||
sig: Regex.exec(text, re) -> Match
|
||||
tip: Like find, but with an already-compiled Regex.
|
||||
order: 6
|
||||
ns: Regex
|
||||
member: exec
|
||||
---
|
||||
|
||||
The compiled-pattern form of <a href="regex-find"><code>Regex.find</code></a>: returns the first <code>Match</code> of the pre-compiled <code>re</code> in <code>text</code>, or <code>null</code>.
|
||||
|
||||
Parameters:
|
||||
- `text` — the string to search
|
||||
- `re` — a compiled `Regex`
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
let re = Regex.compile("#(\\w+)")
|
||||
let m = Regex.exec("a #tag b", re)
|
||||
print(Regex.group(m, 1))
|
||||
}
|
||||
}
|
||||
```
|
||||
27
docs/language/regex/regex-find.md
Normal file
27
docs/language/regex/regex-find.md
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
---
|
||||
id: regex-find
|
||||
name: Regex.find
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.find
|
||||
sig: Regex.find(text, pattern) -> Match
|
||||
tip: The first match of the pattern in the text (a Match, or null if none).
|
||||
order: 5
|
||||
ns: Regex
|
||||
member: find
|
||||
---
|
||||
|
||||
Searches <code>text</code> and returns the first <code>Match</code> — carrying the overall span and every capture group — or <code>null</code> if the pattern does not match. Read groups with <a href="regex-group"><code>Regex.group</code></a> and spans with <a href="regex-start"><code>Regex.start</code></a> / <a href="regex-end"><code>Regex.end</code></a>.
|
||||
|
||||
Parameters:
|
||||
- `text` — the string to search
|
||||
- `pattern` — the regular expression source
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
let m = Regex.find("name: Alice", "(\\w+): (\\w+)")
|
||||
if Regex.ok(m) { print(1) }
|
||||
}
|
||||
}
|
||||
```
|
||||
27
docs/language/regex/regex-group.md
Normal file
27
docs/language/regex/regex-group.md
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
---
|
||||
id: regex-group
|
||||
name: Regex.group
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.group
|
||||
sig: Regex.group(match, n) -> string
|
||||
tip: The text of capture group n (group 0 is the whole match); empty if unset.
|
||||
order: 9
|
||||
ns: Regex
|
||||
member: group
|
||||
---
|
||||
|
||||
Returns the substring captured by group <code>n</code> of <code>match</code> — group <code>0</code> is the whole match, <code>1</code> the first parenthesized group, and so on. Returns an empty string if the group did not participate in the match.
|
||||
|
||||
Parameters:
|
||||
- `match` — a `Match` from `Regex.find` / `Regex.exec` / `Regex.next`
|
||||
- `n` — the group number (0 = whole match)
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
let m = Regex.find("key=value", "(\\w+)=(\\w+)")
|
||||
print(Regex.group(m, 2))
|
||||
}
|
||||
}
|
||||
```
|
||||
26
docs/language/regex/regex-group_count.md
Normal file
26
docs/language/regex/regex-group_count.md
Normal file
|
|
@ -0,0 +1,26 @@
|
|||
---
|
||||
id: regex-group_count
|
||||
name: Regex.group_count
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.group_count
|
||||
sig: Regex.group_count(match) -> int
|
||||
tip: How many capture groups the match's pattern has.
|
||||
order: 10
|
||||
ns: Regex
|
||||
member: group_count
|
||||
---
|
||||
|
||||
Returns the number of capture groups in the pattern behind <code>match</code> (not counting group 0, the whole match). Valid group numbers for <a href="regex-group"><code>Regex.group</code></a> are <code>0</code>…<code>group_count</code>.
|
||||
|
||||
Parameters:
|
||||
- `match` — a `Match` value
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
let m = Regex.find("a-b", "(\\w)-(\\w)")
|
||||
print(Regex.group_count(m))
|
||||
}
|
||||
}
|
||||
```
|
||||
26
docs/language/regex/regex-matches.md
Normal file
26
docs/language/regex/regex-matches.md
Normal file
|
|
@ -0,0 +1,26 @@
|
|||
---
|
||||
id: regex-matches
|
||||
name: Regex.matches
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.matches
|
||||
sig: Regex.matches(text, pattern) -> bool
|
||||
tip: True if the pattern matches anywhere in the text (a search).
|
||||
order: 3
|
||||
ns: Regex
|
||||
member: matches
|
||||
---
|
||||
|
||||
Searches <code>text</code> for <code>pattern</code> and returns whether it matches <em>anywhere</em> (like a search, not a full-string match — anchor with <code>^…$</code> for that). Compiles the pattern each call; for a pattern you use repeatedly, compile once with <a href="regex-compile"><code>Regex.compile</code></a> and use <a href="regex-test"><code>Regex.test</code></a>.
|
||||
|
||||
Parameters:
|
||||
- `text` — the string to search
|
||||
- `pattern` — the regular expression source
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
if Regex.matches("/quit now", "^/(help|quit)\\b") { print(1) }
|
||||
}
|
||||
}
|
||||
```
|
||||
32
docs/language/regex/regex-next.md
Normal file
32
docs/language/regex/regex-next.md
Normal file
|
|
@ -0,0 +1,32 @@
|
|||
---
|
||||
id: regex-next
|
||||
name: Regex.next
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.next
|
||||
sig: Regex.next(text, re, from) -> Match
|
||||
tip: The next match at or after byte offset from — the basis of a find-all loop.
|
||||
order: 7
|
||||
ns: Regex
|
||||
member: next
|
||||
---
|
||||
|
||||
Returns the first match of <code>re</code> in <code>text</code> starting at or after byte offset <code>from</code>, or <code>null</code>. Loop it from the previous match's end (see <a href="regex-end"><code>Regex.end</code></a>) to iterate every match — the find-all idiom.
|
||||
|
||||
Parameters:
|
||||
- `text` — the string to search
|
||||
- `re` — a compiled `Regex`
|
||||
- `from` — the byte offset to resume from
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
let re = Regex.compile("#(\\w+)")
|
||||
var m = Regex.exec("#a #b #c", re)
|
||||
while Regex.ok(m) {
|
||||
print(Regex.group(m, 1))
|
||||
m = Regex.next("#a #b #c", re, Regex.end(m, 0))
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
25
docs/language/regex/regex-ok.md
Normal file
25
docs/language/regex/regex-ok.md
Normal file
|
|
@ -0,0 +1,25 @@
|
|||
---
|
||||
id: regex-ok
|
||||
name: Regex.ok
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.ok
|
||||
sig: Regex.ok(match) -> bool
|
||||
tip: True if the match succeeded (the Match is non-null).
|
||||
order: 13
|
||||
ns: Regex
|
||||
member: ok
|
||||
---
|
||||
|
||||
Returns whether <code>match</code> is a real match rather than <code>null</code> — the readable way to test the result of <a href="regex-find"><code>Regex.find</code></a> / <a href="regex-exec"><code>Regex.exec</code></a> / <a href="regex-next"><code>Regex.next</code></a>, and the loop condition for a find-all.
|
||||
|
||||
Parameters:
|
||||
- `match` — a `Match` value (possibly null)
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
if not Regex.ok(Regex.find("abc", "\\d")) { print(1) }
|
||||
}
|
||||
}
|
||||
```
|
||||
27
docs/language/regex/regex-replace.md
Normal file
27
docs/language/regex/regex-replace.md
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
---
|
||||
id: regex-replace
|
||||
name: Regex.replace
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.replace
|
||||
sig: Regex.replace(text, pattern, replacement) -> string
|
||||
tip: Replace every match; the replacement expands \0..\9 group references.
|
||||
order: 8
|
||||
ns: Regex
|
||||
member: replace
|
||||
---
|
||||
|
||||
Returns a new string with every non-overlapping match of <code>pattern</code> in <code>text</code> replaced by <code>replacement</code>. In the replacement, <code>\0</code> is the whole match, <code>\1</code>…<code>\9</code> are capture groups, and <code>\\</code> is a literal backslash — the natural fit for dialogue/localization token substitution.
|
||||
|
||||
Parameters:
|
||||
- `text` — the string to transform
|
||||
- `pattern` — the regular expression source
|
||||
- `replacement` — the replacement, with `\0`..`\9` group references
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
print(Regex.replace("hi {name}!", "\\{(\\w+)\\}", "[\\1]"))
|
||||
}
|
||||
}
|
||||
```
|
||||
27
docs/language/regex/regex-start.md
Normal file
27
docs/language/regex/regex-start.md
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
---
|
||||
id: regex-start
|
||||
name: Regex.start
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.start
|
||||
sig: Regex.start(match, n) -> int
|
||||
tip: The start byte offset of group n, or -1 if it did not participate.
|
||||
order: 11
|
||||
ns: Regex
|
||||
member: start
|
||||
---
|
||||
|
||||
Returns the byte offset where group <code>n</code> of <code>match</code> begins, or <code>-1</code> if the group did not participate. Group <code>0</code> is the whole match, so <code>Regex.start(m, 0)</code> is where the match begins.
|
||||
|
||||
Parameters:
|
||||
- `match` — a `Match` value
|
||||
- `n` — the group number
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
let m = Regex.find(" hi", "hi")
|
||||
print(Regex.start(m, 0))
|
||||
}
|
||||
}
|
||||
```
|
||||
27
docs/language/regex/regex-test.md
Normal file
27
docs/language/regex/regex-test.md
Normal file
|
|
@ -0,0 +1,27 @@
|
|||
---
|
||||
id: regex-test
|
||||
name: Regex.test
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.test
|
||||
sig: Regex.test(text, re) -> bool
|
||||
tip: Like matches, but reuses an already-compiled Regex.
|
||||
order: 4
|
||||
ns: Regex
|
||||
member: test
|
||||
---
|
||||
|
||||
The compiled-pattern form of <a href="regex-matches"><code>Regex.matches</code></a>: returns whether the pre-compiled <code>re</code> matches anywhere in <code>text</code>. No per-call parsing, so this is the one to call every frame.
|
||||
|
||||
Parameters:
|
||||
- `text` — the string to search
|
||||
- `re` — a compiled `Regex` from `Regex.compile`
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
let re = Regex.compile("[a-z]+")
|
||||
if Regex.test("hello", re) { print(1) }
|
||||
}
|
||||
}
|
||||
```
|
||||
26
docs/language/regex/regex-valid.md
Normal file
26
docs/language/regex/regex-valid.md
Normal file
|
|
@ -0,0 +1,26 @@
|
|||
---
|
||||
id: regex-valid
|
||||
name: Regex.valid
|
||||
category: regex
|
||||
kind: namespace-method
|
||||
tokens: Regex.valid
|
||||
sig: Regex.valid(pattern) -> bool
|
||||
tip: True if the pattern is well-formed (compiles without error).
|
||||
order: 2
|
||||
ns: Regex
|
||||
member: valid
|
||||
---
|
||||
|
||||
A convenience over <a href="regex-compile"><code>Regex.compile</code></a>: <code>true</code> when <code>pattern</code> compiles, <code>false</code> when it is malformed. Use it to validate a user- or mod-supplied pattern before relying on it.
|
||||
|
||||
Parameters:
|
||||
- `pattern` — the regular expression source
|
||||
|
||||
```ludic
|
||||
program Demo {
|
||||
entry {
|
||||
if Regex.valid("(a|b)+") { print(1) }
|
||||
if not Regex.valid("(unterminated") { print(2) }
|
||||
}
|
||||
}
|
||||
```
|
||||
59
examples/library/regex.ludic
Normal file
59
examples/library/regex.ludic
Normal file
|
|
@ -0,0 +1,59 @@
|
|||
# regex.ludic — Regex.* behaviour and the linear-time safety guarantee. Each
|
||||
# assertion that holds prints its number, so a full run prints:
|
||||
# 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18
|
||||
# The engine is a Thompson NFA / Pike VM (see runtime/native/regex.ludic): a
|
||||
# pathological pattern is O(n·m), never exponential — check #17 pins that.
|
||||
program Regex {
|
||||
entry {
|
||||
# --- search (matches) ---
|
||||
if Regex.matches("hello world", "wor.d") { print(1) }
|
||||
if not Regex.matches("hello", "\\d+") { print(2) }
|
||||
|
||||
# --- anchors + alternation + groups (a chat-command validator) ---
|
||||
if Regex.matches("/help", "^/(help|quit)$") { print(3) }
|
||||
if not Regex.matches("/helpme", "^/(help|quit)$") { print(4) }
|
||||
|
||||
# --- character classes / shorthands ---
|
||||
if Regex.matches("abc123", "\\d+") { print(5) }
|
||||
if Regex.matches("user_1@host.io", "^\\w+@\\w+\\.\\w+$") { print(6) }
|
||||
if Regex.matches("x", "[a-z]") { print(7) }
|
||||
if not Regex.matches("5", "[^0-9]") { print(8) }
|
||||
|
||||
# --- greedy vs lazy ---
|
||||
let g = Regex.find("<a><b>", "<(.+)>")
|
||||
if Regex.group(g, 1) == "a><b" { print(9) }
|
||||
let l = Regex.find("<a><b>", "<(.+?)>")
|
||||
if Regex.group(l, 1) == "a" { print(10) }
|
||||
|
||||
# --- bounded repetition {n,m} ---
|
||||
if Regex.matches("aaa", "^a{2,3}$") { print(11) }
|
||||
if not Regex.matches("aaaa", "^a{2,3}$") { print(12) }
|
||||
|
||||
# --- capture groups by number ---
|
||||
let m = Regex.find("name: Alice", "(\\w+): (\\w+)")
|
||||
if Regex.group(m, 1) == "name" and Regex.group(m, 2) == "Alice" { print(13) }
|
||||
|
||||
# --- replace: literal and with a \1 backreference (dialogue tokens) ---
|
||||
if Regex.replace("a1b2c3", "\\d", "#") == "a#b#c#" { print(14) }
|
||||
if Regex.replace("hi {name}!", "\\{(\\w+)\\}", "[\\1]") == "hi [name]!" { print(15) }
|
||||
|
||||
# --- find_all via a compile-once / next loop (extract #hashtags) ---
|
||||
let re = Regex.compile("#(\\w+)")
|
||||
var tags = ""
|
||||
var cur = Regex.exec("#a #bb #ccc", re)
|
||||
while Regex.ok(cur) {
|
||||
tags = tags + Regex.group(cur, 1)
|
||||
cur = Regex.next("#a #bb #ccc", re, Regex.end(cur, 0))
|
||||
}
|
||||
if tags == "abbccc" { print(16) }
|
||||
|
||||
# --- linear time: a classic catastrophic-backtracking pattern still runs
|
||||
# (a naive engine would hang on 40 a's; the NFA is O(n·m)) ---
|
||||
if Regex.matches("aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa", "(a+)+$") { print(17) }
|
||||
|
||||
# --- errors are values: an invalid pattern compiles to null, never crashes ---
|
||||
if not Regex.valid("(unterminated") { print(18) }
|
||||
|
||||
exit(0)
|
||||
}
|
||||
}
|
||||
506
runtime/native/regex.ludic
Normal file
506
runtime/native/regex.ludic
Normal file
|
|
@ -0,0 +1,506 @@
|
|||
# ============================================================================
|
||||
# regex.ludic — a regular-expression engine, written in Ludic.
|
||||
#
|
||||
# PCRE/PECL-compatible syntax over a LINEAR-TIME Thompson NFA (a Pike VM with
|
||||
# capture slots), so a bad pattern from a modder can never cause catastrophic
|
||||
# backtracking — `(a+)+$` on non-matching input is O(n·m), not exponential. The
|
||||
# pattern compiles to a small bytecode program (an unanchored lazy `.*?` prefix
|
||||
# makes a plain search match anywhere); the VM runs every alive thread in
|
||||
# lockstep per input byte, deduped by program counter so work stays bounded.
|
||||
#
|
||||
# Supported: literals, `.`, character classes `[...]` (ranges, negation, and the
|
||||
# \d \w \s \D \W \S shorthands), anchors `^` `$`, alternation `|`, capturing
|
||||
# `(...)` and non-capturing `(?:...)` groups, and the quantifiers `* + ?` and
|
||||
# `{n} {n,} {n,m}` in greedy or lazy (`?`-suffixed) form, plus the common
|
||||
# escapes. Numbered capture groups. Out of scope for a linear engine (and so
|
||||
# unsupported): backreferences and look-around. On the degenerate case of a
|
||||
# nullable subpattern under an unbounded quantifier (e.g. `(a*)*`), match/
|
||||
# capture positions may differ from a backtracking engine like Python's — the
|
||||
# price of the linear-time guarantee.
|
||||
#
|
||||
# ludicc splices this file into any program that mentions `Regex.*` (like the
|
||||
# game runtime in core.ludic); it is a fragment, not a `program`/`module` block.
|
||||
# The Regex.* namespace (emit_call.ludic) aliases each method to the matching
|
||||
# `regex_*` function below.
|
||||
# ============================================================================
|
||||
|
||||
# ---- small helpers ----------------------------------------------------------
|
||||
function rx_slen(s: pointer) -> int { var n = 0; while s[n] != 0 { n = n + 1 }; return n }
|
||||
|
||||
# ---- growable int vector ----------------------------------------------------
|
||||
property IVec { d: words, n: int, cap: int }
|
||||
function iv_new() -> IVec {
|
||||
let v = new IVec
|
||||
v.cap = 16; v.d = words(16); v.n = 0
|
||||
return v
|
||||
}
|
||||
function iv_grow(v: IVec, need: int) -> void {
|
||||
if v.n + need <= v.cap { return }
|
||||
var nc = v.cap * 2
|
||||
while nc < v.n + need { nc = nc * 2 }
|
||||
let nd = words(nc)
|
||||
var i = 0
|
||||
while i < v.n { nd[i] = v.d[i]; i = i + 1 }
|
||||
v.d = nd; v.cap = nc
|
||||
}
|
||||
function iv_push(v: IVec, x: int) -> int {
|
||||
iv_grow(v, 1)
|
||||
let idx = v.n
|
||||
v.d[v.n] = x
|
||||
v.n = v.n + 1
|
||||
return idx
|
||||
}
|
||||
|
||||
# ---- instruction opcodes ----------------------------------------------------
|
||||
const OP_CHAR: int = 1
|
||||
const OP_ANY: int = 2 # any byte except '\n'
|
||||
const OP_CLASS: int = 3
|
||||
const OP_MATCH: int = 4
|
||||
const OP_JMP: int = 5
|
||||
const OP_SPLIT: int = 6
|
||||
const OP_SAVE: int = 7
|
||||
const OP_BOL: int = 8
|
||||
const OP_EOL: int = 9
|
||||
const OP_ANYNL: int = 10 # any byte incl '\n' (the .*? search prefix)
|
||||
|
||||
# ---- AST node types ---------------------------------------------------------
|
||||
const N_LIT: int = 1
|
||||
const N_ANY: int = 2
|
||||
const N_CLASS: int = 3
|
||||
const N_CONCAT: int = 4
|
||||
const N_ALT: int = 5
|
||||
const N_STAR: int = 6
|
||||
const N_PLUS: int = 7
|
||||
const N_QUEST: int = 8
|
||||
const N_REP: int = 9
|
||||
const N_GROUP: int = 10
|
||||
const N_BOL: int = 11
|
||||
const N_EOL: int = 12
|
||||
const N_EMPTY: int = 13
|
||||
|
||||
property RNode {
|
||||
op: int = 0, ch: int = 0, cls: int = 0, lo: int = 0, hi: int = 0,
|
||||
greedy: int = 1, gidx: int = 0 - 1, kids: []RNode
|
||||
}
|
||||
function rx_node(op: int) -> RNode {
|
||||
let n = new RNode
|
||||
n.op = op
|
||||
n.greedy = 1
|
||||
n.gidx = 0 - 1
|
||||
n.kids = new []RNode
|
||||
return n
|
||||
}
|
||||
|
||||
property Prog { code: IVec, cls: IVec, ngroups: int, ok: int }
|
||||
|
||||
# ---- parser state -----------------------------------------------------------
|
||||
var rx_pat: pointer = ""
|
||||
var rx_pos: int = 0
|
||||
var rx_len: int = 0
|
||||
var rx_err: int = 0
|
||||
var rx_ngroup: int = 0
|
||||
var rx_prog: Prog = null
|
||||
|
||||
function rx_peek() -> int { if rx_pos < rx_len { return rx_pat[rx_pos] & 255 }; return 0 - 1 }
|
||||
function rx_peek2() -> int { if rx_pos + 1 < rx_len { return rx_pat[rx_pos + 1] & 255 }; return 0 - 1 }
|
||||
function rx_adv() -> int { let c = rx_peek(); rx_pos = rx_pos + 1; return c }
|
||||
|
||||
# ---- character classes ------------------------------------------------------
|
||||
# a class is 8 i32 words (256 bits) in prog.cls; class k occupies cls[8k .. 8k+8)
|
||||
function rx_class_new() -> int {
|
||||
let idx = rx_prog.cls.n / 8
|
||||
var i = 0
|
||||
while i < 8 { iv_push(rx_prog.cls, 0); i = i + 1 }
|
||||
return idx
|
||||
}
|
||||
function rx_class_set(idx: int, c: int) -> void {
|
||||
let w = idx * 8 + (c >> 5)
|
||||
rx_prog.cls.d[w] = rx_prog.cls.d[w] | (1 << (c & 31))
|
||||
}
|
||||
function rx_class_set_range(idx: int, a: int, b: int) -> void {
|
||||
var c = a
|
||||
while c <= b { rx_class_set(idx, c); c = c + 1 }
|
||||
}
|
||||
function rx_class_negate(idx: int) -> void {
|
||||
var i = 0
|
||||
while i < 8 { let w = idx * 8 + i; rx_prog.cls.d[w] = ~rx_prog.cls.d[w]; i = i + 1 }
|
||||
}
|
||||
function rx_class_unset(idx: int, c: int) -> void {
|
||||
let w = idx * 8 + (c >> 5)
|
||||
rx_prog.cls.d[w] = rx_prog.cls.d[w] & ~(1 << (c & 31))
|
||||
}
|
||||
function rx_class_unset_range(idx: int, a: int, b: int) -> void {
|
||||
var c = a
|
||||
while c <= b { rx_class_unset(idx, c); c = c + 1 }
|
||||
}
|
||||
function rx_class_set_all(idx: int) -> void {
|
||||
var i = 0
|
||||
while i < 8 { rx_prog.cls.d[idx * 8 + i] = 0 - 1; i = i + 1 }
|
||||
}
|
||||
function rx_class_set_word(idx: int) -> void {
|
||||
rx_class_set_range(idx, 48, 57)
|
||||
rx_class_set_range(idx, 65, 90)
|
||||
rx_class_set_range(idx, 97, 122)
|
||||
rx_class_set(idx, 95)
|
||||
}
|
||||
function rx_class_set_ws(idx: int) -> void {
|
||||
rx_class_set(idx, 32); rx_class_set(idx, 9); rx_class_set(idx, 10)
|
||||
rx_class_set(idx, 13); rx_class_set(idx, 12); rx_class_set(idx, 11)
|
||||
}
|
||||
function rx_class_has(prog: Prog, idx: int, c: int) -> bool {
|
||||
let w = prog.cls.d[idx * 8 + (c >> 5)]
|
||||
return (w & (1 << (c & 31))) != 0
|
||||
}
|
||||
function rx_is_digit(c: int) -> bool { return c >= 48 and c <= 57 }
|
||||
function rx_is_word(c: int) -> bool {
|
||||
return (c >= 48 and c <= 57) or (c >= 65 and c <= 90) or (c >= 97 and c <= 122) or c == 95
|
||||
}
|
||||
function rx_is_ws(c: int) -> bool {
|
||||
return c == 32 or c == 9 or c == 10 or c == 13 or c == 12 or c == 11
|
||||
}
|
||||
# OR the set named by \d \D \w \W \s \S into class idx. The negated forms OR in
|
||||
# the complement set bit-by-bit (never via set-all+unset, which would clobber a
|
||||
# previously-set member — e.g. the literal `1` in [1\Da] must survive \D).
|
||||
function rx_class_add_pre(idx: int, kind: int) -> void {
|
||||
if kind == 100 { rx_class_set_range(idx, 48, 57); return } # \d
|
||||
if kind == 119 { rx_class_set_word(idx); return } # \w
|
||||
if kind == 115 { rx_class_set_ws(idx); return } # \s
|
||||
var c = 0
|
||||
while c < 256 {
|
||||
if kind == 68 { if not rx_is_digit(c) { rx_class_set(idx, c) } } # \D
|
||||
else if kind == 87 { if not rx_is_word(c) { rx_class_set(idx, c) } } # \W
|
||||
else if kind == 83 { if not rx_is_ws(c) { rx_class_set(idx, c) } } # \S
|
||||
c = c + 1
|
||||
}
|
||||
}
|
||||
|
||||
# parse a [...] class starting at '['; returns an N_CLASS node
|
||||
function rx_parse_class() -> RNode {
|
||||
rx_adv() # consume '['
|
||||
let idx = rx_class_new()
|
||||
var neg = false
|
||||
if rx_peek() == 94 { neg = true; rx_adv() } # [^ ...]
|
||||
# a ']' as the first char is a literal
|
||||
if rx_peek() == 93 { rx_class_set(idx, 93); rx_adv() }
|
||||
while rx_peek() != 93 and rx_peek() >= 0 {
|
||||
var lo = rx_adv()
|
||||
if lo == 92 { # escape inside class
|
||||
let e = rx_adv()
|
||||
if e == 100 or e == 68 or e == 119 or e == 87 or e == 115 or e == 83 {
|
||||
rx_class_add_pre(idx, e)
|
||||
continue
|
||||
}
|
||||
lo = rx_class_escape_char(e)
|
||||
}
|
||||
# a range a-b (but a trailing '-' before ']' is literal)
|
||||
if rx_peek() == 45 and rx_peek2() != 93 and rx_peek2() >= 0 {
|
||||
rx_adv() # consume '-'
|
||||
var hi = rx_adv()
|
||||
if hi == 92 { hi = rx_class_escape_char(rx_adv()) }
|
||||
rx_class_set_range(idx, lo, hi)
|
||||
} else {
|
||||
rx_class_set(idx, lo)
|
||||
}
|
||||
}
|
||||
if rx_peek() != 93 { rx_err = 1 } else { rx_adv() } # consume ']'
|
||||
if neg { rx_class_negate(idx) }
|
||||
let node = rx_node(N_CLASS)
|
||||
node.cls = idx
|
||||
return node
|
||||
}
|
||||
# map an escaped char inside a class to its byte
|
||||
function rx_class_escape_char(e: int) -> int {
|
||||
if e == 110 { return 10 }
|
||||
if e == 116 { return 9 }
|
||||
if e == 114 { return 13 }
|
||||
if e == 102 { return 12 }
|
||||
if e == 118 { return 11 }
|
||||
if e == 48 { return 0 }
|
||||
return e
|
||||
}
|
||||
|
||||
# ---- escapes outside a class ------------------------------------------------
|
||||
function rx_parse_escape() -> RNode {
|
||||
rx_adv() # consume '\'
|
||||
let e = rx_adv()
|
||||
if e == 100 or e == 68 or e == 119 or e == 87 or e == 115 or e == 83 {
|
||||
let idx = rx_class_new()
|
||||
rx_class_add_pre(idx, e)
|
||||
let node = rx_node(N_CLASS)
|
||||
node.cls = idx
|
||||
return node
|
||||
}
|
||||
if e >= 49 and e <= 57 { rx_err = 1; return rx_node(N_EMPTY) } # backrefs unsupported
|
||||
var ch = e
|
||||
if e == 110 { ch = 10 }
|
||||
else if e == 116 { ch = 9 }
|
||||
else if e == 114 { ch = 13 }
|
||||
else if e == 102 { ch = 12 }
|
||||
else if e == 118 { ch = 11 }
|
||||
else if e == 48 { ch = 0 }
|
||||
else if e < 0 { rx_err = 1; return rx_node(N_EMPTY) }
|
||||
let node = rx_node(N_LIT)
|
||||
node.ch = ch
|
||||
return node
|
||||
}
|
||||
|
||||
# ---- parser -----------------------------------------------------------------
|
||||
function rx_parse_alt() -> RNode {
|
||||
let first = rx_parse_concat()
|
||||
if rx_peek() != 124 { return first }
|
||||
let alt = rx_node(N_ALT)
|
||||
push(alt.kids, first)
|
||||
while rx_peek() == 124 {
|
||||
rx_adv()
|
||||
push(alt.kids, rx_parse_concat())
|
||||
if rx_err == 1 { break }
|
||||
}
|
||||
return alt
|
||||
}
|
||||
function rx_parse_concat() -> RNode {
|
||||
let cat = rx_node(N_CONCAT)
|
||||
while true {
|
||||
let c = rx_peek()
|
||||
if c < 0 or c == 124 or c == 41 { break }
|
||||
push(cat.kids, rx_parse_repeat())
|
||||
if rx_err == 1 { break }
|
||||
}
|
||||
if len(cat.kids) == 1 { return cat.kids[0] }
|
||||
if len(cat.kids) == 0 { return rx_node(N_EMPTY) }
|
||||
return cat
|
||||
}
|
||||
# read the optional lazy '?' after a quantifier; returns greedy flag (0 = lazy)
|
||||
function rx_lazy() -> int {
|
||||
if rx_peek() == 63 { rx_adv(); return 0 }
|
||||
return 1
|
||||
}
|
||||
function rx_parse_repeat() -> RNode {
|
||||
let atom = rx_parse_atom()
|
||||
if rx_err == 1 { return atom }
|
||||
let c = rx_peek()
|
||||
if c == 42 or c == 43 or c == 63 {
|
||||
rx_adv()
|
||||
var op = N_STAR
|
||||
if c == 43 { op = N_PLUS }
|
||||
if c == 63 { op = N_QUEST }
|
||||
let r = rx_node(op)
|
||||
r.greedy = rx_lazy()
|
||||
push(r.kids, atom)
|
||||
return r
|
||||
}
|
||||
if c == 123 { return rx_parse_brace(atom) }
|
||||
return atom
|
||||
}
|
||||
# {n} {n,} {n,m}
|
||||
function rx_parse_brace(atom: RNode) -> RNode {
|
||||
let save = rx_pos
|
||||
rx_adv() # consume '{'
|
||||
var lo = 0
|
||||
var haslo = false
|
||||
while rx_peek() >= 48 and rx_peek() <= 57 { lo = lo * 10 + (rx_adv() - 48); haslo = true }
|
||||
var hi = lo
|
||||
var hasComma = false
|
||||
if rx_peek() == 44 { hasComma = true; rx_adv(); hi = 0 - 1
|
||||
var hashi = false
|
||||
while rx_peek() >= 48 and rx_peek() <= 57 { if not hashi { hi = 0 }; hi = hi * 10 + (rx_adv() - 48); hashi = true }
|
||||
}
|
||||
if rx_peek() != 125 or not haslo { # not a valid brace -> literal '{'
|
||||
rx_pos = save
|
||||
rx_adv()
|
||||
let n = rx_node(N_LIT); n.ch = 123; return n
|
||||
}
|
||||
rx_adv() # consume '}'
|
||||
let r = rx_node(N_REP)
|
||||
r.lo = lo
|
||||
r.hi = hi
|
||||
r.greedy = rx_lazy()
|
||||
push(r.kids, atom)
|
||||
return r
|
||||
}
|
||||
function rx_parse_atom() -> RNode {
|
||||
let c = rx_peek()
|
||||
if c == 40 { # '('
|
||||
rx_adv()
|
||||
var gidx = 0 - 1
|
||||
if rx_peek() == 63 { # (? ...
|
||||
rx_adv()
|
||||
let d = rx_peek()
|
||||
if d == 58 { rx_adv() } # (?: non-capturing
|
||||
else if d == 80 { rx_adv(); rx_skip_name() } # (?P<name> (captured, name ignored for now)
|
||||
else if d == 60 { rx_skip_name() } # (?<name>
|
||||
else { rx_err = 1; return rx_node(N_EMPTY) }
|
||||
} else {
|
||||
rx_ngroup = rx_ngroup + 1
|
||||
gidx = rx_ngroup
|
||||
}
|
||||
let inner = rx_parse_alt()
|
||||
if rx_peek() != 41 { rx_err = 1; return rx_node(N_EMPTY) }
|
||||
rx_adv() # consume ')'
|
||||
let g = rx_node(N_GROUP)
|
||||
g.gidx = gidx
|
||||
push(g.kids, inner)
|
||||
return g
|
||||
}
|
||||
if c == 91 { return rx_parse_class() }
|
||||
if c == 46 { rx_adv(); return rx_node(N_ANY) }
|
||||
if c == 94 { rx_adv(); return rx_node(N_BOL) }
|
||||
if c == 36 { rx_adv(); return rx_node(N_EOL) }
|
||||
if c == 92 { return rx_parse_escape() }
|
||||
if c == 42 or c == 43 or c == 63 or c == 41 { rx_err = 1; return rx_node(N_EMPTY) }
|
||||
if c < 0 { rx_err = 1; return rx_node(N_EMPTY) }
|
||||
rx_adv()
|
||||
let n = rx_node(N_LIT)
|
||||
n.ch = c
|
||||
return n
|
||||
}
|
||||
# skip a (?P<name> / (?<name> group name up to '>'; leaves a capturing group
|
||||
function rx_skip_name() -> void {
|
||||
if rx_peek() == 80 { rx_adv() } # already consumed by caller in (?P case? guard
|
||||
if rx_peek() == 60 { rx_adv() } # consume '<'
|
||||
while rx_peek() != 62 and rx_peek() >= 0 { rx_adv() }
|
||||
if rx_peek() == 62 { rx_adv() } # consume '>'
|
||||
rx_ngroup = rx_ngroup + 1
|
||||
# note: the caller set gidx = -1; fix it up to a real capture index
|
||||
rx_named_gidx = rx_ngroup
|
||||
}
|
||||
var rx_named_gidx: int = 0
|
||||
|
||||
# ---- compile AST -> program -------------------------------------------------
|
||||
function pg_emit(op: int, a: int, b: int) -> int {
|
||||
let pc = rx_prog.code.n / 3
|
||||
iv_push(rx_prog.code, op)
|
||||
iv_push(rx_prog.code, a)
|
||||
iv_push(rx_prog.code, b)
|
||||
return pc
|
||||
}
|
||||
function pg_set_a(pc: int, a: int) -> void { rx_prog.code.d[3 * pc + 1] = a }
|
||||
function pg_set_b(pc: int, b: int) -> void { rx_prog.code.d[3 * pc + 2] = b }
|
||||
function pg_pc() -> int { return rx_prog.code.n / 3 }
|
||||
|
||||
function rx_compile(node: RNode) -> void {
|
||||
let op = node.op
|
||||
if op == N_EMPTY { return }
|
||||
if op == N_LIT { pg_emit(OP_CHAR, node.ch, 0); return }
|
||||
if op == N_ANY { pg_emit(OP_ANY, 0, 0); return }
|
||||
if op == N_CLASS { pg_emit(OP_CLASS, node.cls, 0); return }
|
||||
if op == N_BOL { pg_emit(OP_BOL, 0, 0); return }
|
||||
if op == N_EOL { pg_emit(OP_EOL, 0, 0); return }
|
||||
if op == N_CONCAT {
|
||||
var i = 0
|
||||
while i < len(node.kids) { rx_compile(node.kids[i]); i = i + 1 }
|
||||
return
|
||||
}
|
||||
if op == N_GROUP {
|
||||
if node.gidx >= 0 {
|
||||
pg_emit(OP_SAVE, 2 * node.gidx, 0)
|
||||
rx_compile(node.kids[0])
|
||||
pg_emit(OP_SAVE, 2 * node.gidx + 1, 0)
|
||||
} else {
|
||||
rx_compile(node.kids[0])
|
||||
}
|
||||
return
|
||||
}
|
||||
if op == N_ALT {
|
||||
let jmps = iv_new()
|
||||
var i = 0
|
||||
while i < len(node.kids) {
|
||||
if i < len(node.kids) - 1 {
|
||||
let sp = pg_emit(OP_SPLIT, 0, 0)
|
||||
pg_set_a(sp, pg_pc())
|
||||
rx_compile(node.kids[i])
|
||||
let j = pg_emit(OP_JMP, 0, 0)
|
||||
iv_push(jmps, j)
|
||||
pg_set_b(sp, pg_pc())
|
||||
} else {
|
||||
rx_compile(node.kids[i])
|
||||
}
|
||||
i = i + 1
|
||||
}
|
||||
let end = pg_pc()
|
||||
i = 0
|
||||
while i < jmps.n { pg_set_a(jmps.d[i], end); i = i + 1 }
|
||||
return
|
||||
}
|
||||
if op == N_STAR {
|
||||
let l1 = pg_pc()
|
||||
let sp = pg_emit(OP_SPLIT, 0, 0)
|
||||
let l2 = pg_pc()
|
||||
rx_compile(node.kids[0])
|
||||
pg_emit(OP_JMP, l1, 0)
|
||||
let l3 = pg_pc()
|
||||
if node.greedy == 1 { pg_set_a(sp, l2); pg_set_b(sp, l3) }
|
||||
else { pg_set_a(sp, l3); pg_set_b(sp, l2) }
|
||||
return
|
||||
}
|
||||
if op == N_PLUS {
|
||||
let l1 = pg_pc()
|
||||
rx_compile(node.kids[0])
|
||||
let sp = pg_emit(OP_SPLIT, 0, 0)
|
||||
let l3 = pg_pc()
|
||||
if node.greedy == 1 { pg_set_a(sp, l1); pg_set_b(sp, l3) }
|
||||
else { pg_set_a(sp, l3); pg_set_b(sp, l1) }
|
||||
return
|
||||
}
|
||||
if op == N_QUEST {
|
||||
let sp = pg_emit(OP_SPLIT, 0, 0)
|
||||
let l2 = pg_pc()
|
||||
rx_compile(node.kids[0])
|
||||
let l3 = pg_pc()
|
||||
if node.greedy == 1 { pg_set_a(sp, l2); pg_set_b(sp, l3) }
|
||||
else { pg_set_a(sp, l3); pg_set_b(sp, l2) }
|
||||
return
|
||||
}
|
||||
if op == N_REP {
|
||||
let kid = node.kids[0]
|
||||
var i = 0
|
||||
while i < node.lo { rx_compile(kid); i = i + 1 }
|
||||
if node.hi < 0 {
|
||||
let st = rx_node(N_STAR)
|
||||
st.greedy = node.greedy
|
||||
push(st.kids, kid)
|
||||
rx_compile(st)
|
||||
} else {
|
||||
i = 0
|
||||
while i < node.hi - node.lo {
|
||||
let q = rx_node(N_QUEST)
|
||||
q.greedy = node.greedy
|
||||
push(q.kids, kid)
|
||||
rx_compile(q)
|
||||
i = i + 1
|
||||
}
|
||||
}
|
||||
return
|
||||
}
|
||||
}
|
||||
|
||||
# regex_compile(pattern) -> Prog (null on a syntax error)
|
||||
function regex_compile(pattern: pointer) -> Prog {
|
||||
let p = new Prog
|
||||
p.code = iv_new()
|
||||
p.cls = iv_new()
|
||||
p.ngroups = 0
|
||||
p.ok = 1
|
||||
rx_prog = p
|
||||
rx_pat = pattern
|
||||
rx_pos = 0
|
||||
rx_len = rx_slen(pattern)
|
||||
rx_err = 0
|
||||
rx_ngroup = 0
|
||||
let root = rx_parse_alt()
|
||||
if rx_err == 1 or rx_pos != rx_len { return null }
|
||||
p.ngroups = rx_ngroup
|
||||
# unanchored lazy .*? prefix so a match may start at any position
|
||||
let sp0 = pg_emit(OP_SPLIT, 0, 0)
|
||||
let consume = pg_pc()
|
||||
pg_emit(OP_ANYNL, 0, 0)
|
||||
pg_emit(OP_JMP, sp0, 0)
|
||||
let body = pg_pc()
|
||||
pg_set_a(sp0, body)
|
||||
pg_set_b(sp0, consume)
|
||||
pg_emit(OP_SAVE, 0, 0)
|
||||
rx_compile(root)
|
||||
pg_emit(OP_SAVE, 1, 0)
|
||||
pg_emit(OP_MATCH, 0, 0)
|
||||
return p
|
||||
}
|
||||
|
||||
211
runtime/native/regex_vm.ludic
Normal file
211
runtime/native/regex_vm.ludic
Normal file
|
|
@ -0,0 +1,211 @@
|
|||
# ============================================================================
|
||||
# regex_vm.ludic — the Pike VM (Thompson NFA execution) and the public Regex.*
|
||||
# API for the engine compiled in regex.ludic. Splits the "run + query" concern
|
||||
# out of regex.ludic's "parse + compile". Spliced together (regex.ludic imports
|
||||
# this file), so the two share the Prog/RNode model and OP_* opcodes.
|
||||
# ============================================================================
|
||||
|
||||
# ---- Pike VM ----------------------------------------------------------------
|
||||
property TList { pc: words, caps: words, n: int }
|
||||
var rx_seen: words = null
|
||||
var rx_gen: int = 0
|
||||
var rx_nc: int = 0
|
||||
var rx_code: words = null
|
||||
|
||||
function rx_add(list: TList, pc: int, caps: words, sp: int, s: pointer, slen: int) -> void {
|
||||
if rx_seen[pc] == rx_gen { return }
|
||||
rx_seen[pc] = rx_gen
|
||||
let op = rx_code[3 * pc]
|
||||
if op == OP_JMP { rx_add(list, rx_code[3 * pc + 1], caps, sp, s, slen); return }
|
||||
if op == OP_SPLIT {
|
||||
rx_add(list, rx_code[3 * pc + 1], caps, sp, s, slen)
|
||||
rx_add(list, rx_code[3 * pc + 2], caps, sp, s, slen)
|
||||
return
|
||||
}
|
||||
if op == OP_SAVE {
|
||||
let slot = rx_code[3 * pc + 1]
|
||||
let old = caps[slot]
|
||||
caps[slot] = sp
|
||||
rx_add(list, pc + 1, caps, sp, s, slen)
|
||||
caps[slot] = old
|
||||
return
|
||||
}
|
||||
if op == OP_BOL {
|
||||
if sp == 0 { rx_add(list, pc + 1, caps, sp, s, slen) }
|
||||
return
|
||||
}
|
||||
if op == OP_EOL {
|
||||
if sp == slen { rx_add(list, pc + 1, caps, sp, s, slen) }
|
||||
else if sp == slen - 1 and (s[sp] & 255) == 10 { rx_add(list, pc + 1, caps, sp, s, slen) }
|
||||
return
|
||||
}
|
||||
# a leaf that consumes (CHAR/ANY/ANYNL/CLASS) or MATCH: record it
|
||||
let t = list.n
|
||||
list.pc[t] = pc
|
||||
var i = 0
|
||||
while i < rx_nc { list.caps[t * rx_nc + i] = caps[i]; i = i + 1 }
|
||||
list.n = list.n + 1
|
||||
}
|
||||
|
||||
# run prog over s (length slen) from startpos; returns caps words or null
|
||||
function rx_run(prog: Prog, s: pointer, slen: int, startpos: int) -> words {
|
||||
rx_code = prog.code.d
|
||||
rx_nc = 2 * (prog.ngroups + 1)
|
||||
let ncode = prog.code.n / 3
|
||||
rx_seen = words(ncode)
|
||||
var i = 0
|
||||
while i < ncode { rx_seen[i] = 0; i = i + 1 }
|
||||
rx_gen = 0
|
||||
var clist = new TList
|
||||
clist.pc = words(ncode); clist.caps = words(ncode * rx_nc); clist.n = 0
|
||||
var nlist = new TList
|
||||
nlist.pc = words(ncode); nlist.caps = words(ncode * rx_nc); nlist.n = 0
|
||||
let wcaps = words(rx_nc)
|
||||
var matched: words = null
|
||||
|
||||
rx_gen = rx_gen + 1
|
||||
i = 0
|
||||
while i < rx_nc { wcaps[i] = 0 - 1; i = i + 1 }
|
||||
rx_add(clist, 0, wcaps, startpos, s, slen)
|
||||
|
||||
var sp = startpos
|
||||
while true {
|
||||
if clist.n == 0 { break }
|
||||
var c = 0 - 1
|
||||
if sp < slen { c = s[sp] & 255 }
|
||||
nlist.n = 0
|
||||
rx_gen = rx_gen + 1
|
||||
var ti = 0
|
||||
var stop = false
|
||||
while ti < clist.n and not stop {
|
||||
let pc = clist.pc[ti]
|
||||
var k = 0
|
||||
while k < rx_nc { wcaps[k] = clist.caps[ti * rx_nc + k]; k = k + 1 }
|
||||
let op = rx_code[3 * pc]
|
||||
if op == OP_CHAR {
|
||||
if c >= 0 and c == rx_code[3 * pc + 1] { rx_add(nlist, pc + 1, wcaps, sp + 1, s, slen) }
|
||||
} else if op == OP_ANY {
|
||||
if c >= 0 and c != 10 { rx_add(nlist, pc + 1, wcaps, sp + 1, s, slen) }
|
||||
} else if op == OP_ANYNL {
|
||||
if c >= 0 { rx_add(nlist, pc + 1, wcaps, sp + 1, s, slen) }
|
||||
} else if op == OP_CLASS {
|
||||
if c >= 0 and rx_class_has(prog, rx_code[3 * pc + 1], c) { rx_add(nlist, pc + 1, wcaps, sp + 1, s, slen) }
|
||||
} else if op == OP_MATCH {
|
||||
if matched == null { matched = words(rx_nc) }
|
||||
k = 0
|
||||
while k < rx_nc { matched[k] = wcaps[k]; k = k + 1 }
|
||||
stop = true
|
||||
}
|
||||
ti = ti + 1
|
||||
}
|
||||
let tmp = clist; clist = nlist; nlist = tmp
|
||||
if sp >= slen { break }
|
||||
sp = sp + 1
|
||||
}
|
||||
return matched
|
||||
}
|
||||
|
||||
# ---- public API -------------------------------------------------------------
|
||||
property Match { str: pointer = null, ng: int = 0, caps: words = null }
|
||||
|
||||
function regex_matches(str: pointer, pattern: pointer) -> bool {
|
||||
let p = regex_compile(pattern)
|
||||
if p == null { return false }
|
||||
return rx_run(p, str, rx_slen(str), 0) != null
|
||||
}
|
||||
function regex_test(str: pointer, re: Prog) -> bool {
|
||||
if re == null { return false }
|
||||
return rx_run(re, str, rx_slen(str), 0) != null
|
||||
}
|
||||
function regex_exec(str: pointer, re: Prog) -> Match {
|
||||
if re == null { return null }
|
||||
let caps = rx_run(re, str, rx_slen(str), 0)
|
||||
if caps == null { return null }
|
||||
let m = new Match
|
||||
m.str = str; m.ng = re.ngroups; m.caps = caps
|
||||
return m
|
||||
}
|
||||
function regex_find(str: pointer, pattern: pointer) -> Match {
|
||||
let p = regex_compile(pattern)
|
||||
if p == null { return null }
|
||||
return regex_exec(str, p)
|
||||
}
|
||||
function regex_next(str: pointer, re: Prog, from: int) -> Match {
|
||||
if re == null { return null }
|
||||
let n = rx_slen(str)
|
||||
if from > n { return null }
|
||||
let caps = rx_run(re, str, n, from)
|
||||
if caps == null { return null }
|
||||
let m = new Match
|
||||
m.str = str; m.ng = re.ngroups; m.caps = caps
|
||||
return m
|
||||
}
|
||||
function regex_ok(m: Match) -> bool { return m != null }
|
||||
function regex_group_count(m: Match) -> int { if m == null { return 0 }; return m.ng }
|
||||
function regex_start(m: Match, n: int) -> int {
|
||||
if m == null or n < 0 or n > m.ng { return 0 - 1 }
|
||||
return m.caps[2 * n]
|
||||
}
|
||||
function regex_end(m: Match, n: int) -> int {
|
||||
if m == null or n < 0 or n > m.ng { return 0 - 1 }
|
||||
return m.caps[2 * n + 1]
|
||||
}
|
||||
function regex_group(m: Match, n: int) -> pointer {
|
||||
if m == null or n < 0 or n > m.ng { return "" }
|
||||
let a = m.caps[2 * n]
|
||||
let b = m.caps[2 * n + 1]
|
||||
if a < 0 or b < 0 { return "" }
|
||||
let out = bytes(b - a + 1)
|
||||
var i = 0
|
||||
while i < b - a { out[i] = m.str[a + i]; i = i + 1 }
|
||||
out[b - a] = 0
|
||||
return out
|
||||
}
|
||||
function regex_valid(pattern: pointer) -> bool { return regex_compile(pattern) != null }
|
||||
|
||||
# regex_replace(str, pattern, repl): replace all non-overlapping matches.
|
||||
# repl expands \0..\9 (groups; \0 = whole match) and \\ (a literal backslash).
|
||||
function rx_append(out: IVec, s: pointer, a: int, b: int) -> void {
|
||||
var i = a
|
||||
while i < b { iv_push(out, s[i] & 255); i = i + 1 }
|
||||
}
|
||||
function regex_replace(str: pointer, pattern: pointer, repl: pointer) -> pointer {
|
||||
let p = regex_compile(pattern)
|
||||
if p == null { return str }
|
||||
let n = rx_slen(str)
|
||||
let rn = rx_slen(repl)
|
||||
let out = iv_new() # bytes accumulated as ints
|
||||
var pos = 0
|
||||
var prev = 0
|
||||
while pos <= n {
|
||||
let caps = rx_run(p, str, n, pos)
|
||||
if caps == null { break }
|
||||
let ms = caps[0]
|
||||
let me = caps[1]
|
||||
rx_append(out, str, prev, ms) # text before the match
|
||||
# expand replacement
|
||||
var i = 0
|
||||
while i < rn {
|
||||
if repl[i] == 92 and i + 1 < rn {
|
||||
let d = repl[i + 1] & 255
|
||||
if d >= 48 and d <= 57 {
|
||||
let g = d - 48
|
||||
if g <= p.ngroups { let a = caps[2 * g]; let b = caps[2 * g + 1]; if a >= 0 and b >= 0 { rx_append(out, str, a, b) } }
|
||||
i = i + 2
|
||||
} else if d == 92 { iv_push(out, 92); i = i + 2 }
|
||||
else { iv_push(out, d); i = i + 2 }
|
||||
} else {
|
||||
iv_push(out, repl[i] & 255); i = i + 1
|
||||
}
|
||||
}
|
||||
prev = me
|
||||
if me > pos { pos = me } else { if me < n { iv_push(out, str[me] & 255) }; prev = me + 1; pos = me + 1 }
|
||||
}
|
||||
rx_append(out, str, prev, n) # trailing text
|
||||
let res = bytes(out.n + 1)
|
||||
var i = 0
|
||||
while i < out.n { res[i] = out.d[i]; i = i + 1 }
|
||||
res[out.n] = 0
|
||||
return res
|
||||
}
|
||||
|
||||
|
|
@ -185,6 +185,24 @@ function emit_ns_call(ns: pointer, meth: pointer, e: Node) -> Val {
|
|||
if (meth == "write") { bare = "save" }
|
||||
if (meth == "read") { bare = "load" }
|
||||
}
|
||||
# Regex.* -> the regex_* engine functions (spliced from runtime/native/regex*.ludic
|
||||
# when a program mentions Regex.*). Each is a plain alias; the engine functions
|
||||
# are ordinary Ludic, so the generic call path resolves them to @fn_regex_*.
|
||||
if (ns == "Regex") {
|
||||
if (meth == "compile") { bare = "regex_compile"; push(labels, "pattern") }
|
||||
if (meth == "valid") { bare = "regex_valid"; push(labels, "pattern") }
|
||||
if (meth == "matches") { bare = "regex_matches"; push(labels, "text"); push(labels, "pattern") }
|
||||
if (meth == "test") { bare = "regex_test"; push(labels, "text"); push(labels, "re") }
|
||||
if (meth == "find") { bare = "regex_find"; push(labels, "text"); push(labels, "pattern") }
|
||||
if (meth == "exec") { bare = "regex_exec"; push(labels, "text"); push(labels, "re") }
|
||||
if (meth == "next") { bare = "regex_next"; push(labels, "text"); push(labels, "re"); push(labels, "from") }
|
||||
if (meth == "replace") { bare = "regex_replace"; push(labels, "text"); push(labels, "pattern"); push(labels, "replacement") }
|
||||
if (meth == "group") { bare = "regex_group"; push(labels, "match"); push(labels, "n") }
|
||||
if (meth == "group_count") { bare = "regex_group_count"; push(labels, "match") }
|
||||
if (meth == "start") { bare = "regex_start"; push(labels, "match"); push(labels, "n") }
|
||||
if (meth == "end") { bare = "regex_end"; push(labels, "match"); push(labels, "n") }
|
||||
if (meth == "ok") { bare = "regex_ok"; push(labels, "match") }
|
||||
}
|
||||
if (bare == null) { perr(`unknown builtin {ns}.{meth}`) }
|
||||
reorder_named(e, labels)
|
||||
let id = node(E_ID); id.s = bare; e.a = id
|
||||
|
|
|
|||
|
|
@ -160,7 +160,9 @@ function p_primary() -> Node {
|
|||
function p_postfix() -> Node {
|
||||
var e = p_primary()
|
||||
while true {
|
||||
if is_op(".") { pi = pi + 1; let m = node(E_MEMBER); m.a = e; m.s = eat_id(); e = m }
|
||||
if is_op(".") { pi = pi + 1; let m = node(E_MEMBER); m.a = e; m.s = eat_id(); e = m
|
||||
if e.a.kind == E_ID and e.a.s == "Regex" { g_uses_regex = true } # splice the regex runtime on demand
|
||||
}
|
||||
else { if is_op("[") { pi = pi + 1; let lo = expr()
|
||||
if is_op("..") { pi = pi + 1; let sl = node(E_SLICE); sl.a = e; sl.b = lo; sl.c = expr(); eat_op("]"); e = sl } # s[a..b] substring
|
||||
else { let ix = node(E_INDEX); ix.a = e; ix.b = lo; eat_op("]"); e = ix } }
|
||||
|
|
@ -371,6 +373,7 @@ function path_join(dir: pointer, rel: pointer) -> pointer {
|
|||
|
||||
var loaded_paths: []pointer
|
||||
var cur_dir: pointer
|
||||
var g_uses_regex: bool = false # a program mentioned Regex.* -> splice the regex runtime
|
||||
|
||||
function already_loaded(full: pointer) -> bool {
|
||||
var i = 0
|
||||
|
|
@ -507,11 +510,21 @@ function do_import(rel: pointer) -> void {
|
|||
# a game (has systems/components) links the Ludic runtime; auto-splice it the
|
||||
# way the C compiler does. Tools (a `main` block, no ECS) get nothing.
|
||||
function maybe_splice_runtime() -> void {
|
||||
if not has_ecs() { return }
|
||||
let saved = cur_dir
|
||||
cur_dir = ""
|
||||
do_import("runtime/native/core.ludic")
|
||||
cur_dir = saved
|
||||
# a game (has systems/components) links the Ludic runtime.
|
||||
if has_ecs() {
|
||||
cur_dir = ""
|
||||
do_import("runtime/native/core.ludic")
|
||||
cur_dir = saved
|
||||
}
|
||||
# any program that uses Regex.* gets the regex engine spliced in (it is
|
||||
# self-contained — only compiler intrinsics — so it works in a plain tool too).
|
||||
if g_uses_regex {
|
||||
cur_dir = ""
|
||||
do_import("runtime/native/regex.ludic")
|
||||
do_import("runtime/native/regex_vm.ludic")
|
||||
cur_dir = saved
|
||||
}
|
||||
}
|
||||
|
||||
function parse_program() -> void {
|
||||
|
|
@ -529,6 +542,7 @@ function parse_program() -> void {
|
|||
g_events = new []Node
|
||||
g_onlisten = new []Node
|
||||
g_toggled_layers = new []pointer
|
||||
g_uses_regex = false
|
||||
loaded_paths = new []pointer
|
||||
skipnl()
|
||||
g_game_name = "Ludic"
|
||||
|
|
|
|||
26928
selfhost/ludicc.seed.ll
26928
selfhost/ludicc.seed.ll
File diff suppressed because it is too large
Load diff
|
|
@ -486,5 +486,20 @@
|
|||
"mime": [
|
||||
"mime-of",
|
||||
"mime-sniff"
|
||||
],
|
||||
"regex": [
|
||||
"regex-compile",
|
||||
"regex-valid",
|
||||
"regex-matches",
|
||||
"regex-test",
|
||||
"regex-find",
|
||||
"regex-exec",
|
||||
"regex-next",
|
||||
"regex-replace",
|
||||
"regex-group",
|
||||
"regex-group_count",
|
||||
"regex-start",
|
||||
"regex-end",
|
||||
"regex-ok"
|
||||
]
|
||||
}
|
||||
|
|
|
|||
|
|
@ -102,6 +102,7 @@ function cmd_test() -> int {
|
|||
feat_case("library/crypto", "", "1 2 3 4 5 6 7 8 9", "crypto.ludic (Crypto SHA-256/HMAC/base64 KAT + CSPRNG shape)")
|
||||
feat_case("library/uuid", "", "1 2 3 4 5 6 7 8 9 10", "uuid.ludic (Uuid v4/v7 format, version/variant, parse/equals)")
|
||||
feat_case("library/noise", "", "1 2 3 4 5 6 7 8 9 10 11", "noise.ludic (Noise value/perlin/simplex/fbm/cellular determinism + range)")
|
||||
feat_case("library/regex", "", "1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18", "regex.ludic (Regex match/find/groups/classes/quantifiers/replace + linear-time safety)")
|
||||
feat_case("library/logging", "", "0 5 2 1", "logging.ludic (Log levels, set_level/level threshold, structured fields)")
|
||||
# Os known-folders/arch and Fs.list read the BSD utsname/dirent layout, so
|
||||
# their asserted values are macOS-specific; skip off Darwin (see is_darwin).
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue