feat(jobs): real OS threads - Job.parallel_for, fn name, thread-safe Sync

- `fn name` names a top-level function as a value (E_FNREF, lowers to @fn_<name>); the worker
  entry point for Job.parallel_for, which checks it takes (int, pointer-like) and returns void.
- runtime/native/threads.ll (pthreads) and threads_win.ll (Win32 SRWLOCK/CONDITION_VARIABLE): a
  pool of one worker per core but one, parked between batches; every thread claims chunks by
  compare-and-swap. Linked only into programs that use Job/Promise/Sync, by `ludicc -o`,
  `ludic build` and the test suite's build helper.
- Sync.* is real: native mutexes, atomics as cmpxchg retry loops (neither clang takes atomicrw,
  the PC's rejects seq_consistent), mutex-guarded channels, Sync.cpu_count from the OS.
- spawn/despawn on a pool thread stop the program with a located panic.
- examples/library/threads.ludic and its test; docs for fn, Job.parallel_for, Job.is_worker.
- Reseeded (bootstrap-cfree: out.ll == seed.ll). 141/141 on macOS; jobs, threads and the guard
  pass on Windows from the reseeded Windows seed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Orkun ÇAKILKAYA 2026-09-15 13:52:22 +03:00
parent f5d1a62ccf
commit c10abd9f9f
22 changed files with 57464 additions and 55794 deletions

View file

@ -4,4 +4,4 @@ title: Job
order: 37
---
Background work that stays out of the frame. A <code>Job</code> is a future — a handle to a result that lands later. Kick one off with <code>Job.run</code> (a background compute that advances a little each <code>Job.pump</code> and finishes after enough frames, so heavy work never hitches) or <code>Job.defer</code> (a future you resolve yourself with <code>Job.fulfill</code> / <code>Job.fail</code>). Poll it with <code>done</code> / <code>ok</code> / <code>failed</code> / <code>cancelled</code>, read <code>result</code> / <code>error</code>, and always collect on the main thread — a Job must never touch the ECS world directly. The scheduler is deterministic and cooperative, so the same jobs and the same budget reproduce byte-for-byte, every run and every target. Arguments are positional. Spliced in only when a program mentions <code>Job.*</code>.
Background work that stays out of the frame. A <code>Job</code> is a future — a handle to a result that lands later. Kick one off with <code>Job.run</code> (a background compute that advances a little each <code>Job.pump</code> and finishes after enough frames, so heavy work never hitches) or <code>Job.defer</code> (a future you resolve yourself with <code>Job.fulfill</code> / <code>Job.fail</code>). Poll it with <code>done</code> / <code>ok</code> / <code>failed</code> / <code>cancelled</code>, read <code>result</code> / <code>error</code>, and always collect on the main thread — a Job must never touch the ECS world directly. The scheduler is deterministic and cooperative, so the same jobs and the same budget reproduce byte-for-byte, every run and every target. For work that should use every core now, <code>Job.parallel_for(count, fn work, ctx)</code> runs <code>work(i, ctx)</code> across a pool of real OS threads and returns when all of it is done; a worker computes on what it was handed and never changes the world. Arguments are positional. Spliced in only when a program mentions <code>Job.*</code>.

View file

@ -0,0 +1,19 @@
---
id: job-is_worker
name: Job.is_worker
category: job
kind: namespace-method
tokens: Job.is_worker
sig: Job.is_worker() -> bool
tip: True on a Job.parallel_for pool thread, false on the main thread.
order: 17
ns: Job
member: is_worker
---
True when the code is running on a `Job.parallel_for` pool thread, false on the main thread (and on
the calling thread while it takes its own share of a batch).
```ludic
if not Job.is_worker() { print("main thread") }
```

View file

@ -0,0 +1,34 @@
---
id: job-parallel_for
name: Job.parallel_for
category: job
kind: namespace-method
tokens: Job.parallel_for
sig: Job.parallel_for(count, work, ctx) -> void
tip: Run work(i, ctx) for every i in [0, count) across all cores; returns when every call is done.
order: 16
ns: Job
member: parallel_for
---
Run `work(i, ctx)` for every `i` in `[0, count)` on real OS threads, and return when every call has
finished. The indices are split into chunks shared by a worker pool (one thread per core but one,
started on first use) and the calling thread, so each index runs exactly once, in no particular order.
`work` is a top-level function taking `(i: int, ctx)`, named with `fn`; `ctx` is whatever it computes
on (`words`, `bytes` or a `pointer`). A worker computes on what it was handed and writes only its own
index's results: it must not `spawn`, `despawn`, `push` onto a list another thread can see, or use
Http or Audio. `spawn` and `despawn` on a worker stop the program with a located message. Shared
counters go through `Sync.add`, shared totals behind a `Sync.mutex`.
```ludic
program Squares {
function square(i: int, out: words) -> void { out[i] = i * i }
entry {
let sq = words(20000)
Job.parallel_for(20000, fn square, sq)
print(sq[141])
}
}
```

View file

@ -0,0 +1,24 @@
---
id: kw-fn
name: fn
category: structure
kind: keyword
tokens: fn
sig: fn name
tip: A top-level function named as a value - the entry point handed to Job.parallel_for.
order: 9
---
<code>fn name</code> names a top-level function as a value, so it can be handed to something that calls it later - today, <code>Job.parallel_for</code>, which runs it on worker threads. It is a plain function reference: no closure and nothing captured, so everything the function needs is passed to it (for <code>Job.parallel_for</code>, through its <code>ctx</code> argument). A worker function takes <code>(i: int, ctx)</code> and returns <code>void</code>; the compiler rejects any other shape.
```ludic
program Squares {
function square(i: int, out: words) -> void { out[i] = i * i }
entry {
let sq = words(1000)
Job.parallel_for(1000, fn square, sq)
print(sq[12])
}
}
```

View file

@ -5,13 +5,14 @@ category: sync
kind: namespace-method
tokens: Sync.cpu_count
sig: Sync.cpu_count() -> int
tip: Worker lanes available to the scheduler.
tip: The machine's logical cores (1 to 64); Job.parallel_for uses one thread per core.
order: 14
ns: Sync
member: cpu_count
---
Worker lanes available to the scheduler.
The machine's logical cores, from 1 to 64. `Job.parallel_for` runs on that many threads: a pool of
one per core but one, plus the caller.
```ludic
let lanes = Sync.cpu_count()