Within a pass a row writes only its own still-masked texels and reads only neighbours already let go, which no row writes that pass, so the result is the single-threaded one. The worker takes a DilateJob of plain buffers and allocates nothing. tex_dilate_bytes (safe_api) and examples/rendering/dilate.ludic, which checks it against the old loop on RGB and RGBA atlases of sizes that do not divide (DILATE OK). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
8 lines
609 B
Markdown
8 lines
609 B
Markdown
bump: patch
|
|
type: performance
|
|
**Cut-out edge padding runs on every core.** `tex_dilate`'s passes hand their rows, sixteen at a time, to
|
|
`Job.parallel_for`: within a pass a row writes only its own still-masked texels and reads only
|
|
neighbours the mask already let go, so the bytes are the ones the single-threaded loop made. The worker
|
|
is handed plain buffers in a `DilateJob` and makes nothing. `tex_dilate_bytes` is the slice-taking
|
|
form (safe_api.ludic), and `examples/rendering/dilate.ludic` holds the result against the old loop
|
|
(prints DILATE OK). It was 206 ms of the main thread in a Maroon Lake boot.
|