Within a pass a row writes only its own still-masked texels and reads only neighbours already let go, which no row writes that pass, so the result is the single-threaded one. The worker takes a DilateJob of plain buffers and allocates nothing. tex_dilate_bytes (safe_api) and examples/rendering/dilate.ludic, which checks it against the old loop on RGB and RGBA atlases of sizes that do not divide (DILATE OK). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
609 B
609 B
bump: patch
type: performance
Cut-out edge padding runs on every core. tex_dilate's passes hand their rows, sixteen at a time, to
Job.parallel_for: within a pass a row writes only its own still-masked texels and reads only
neighbours the mask already let go, so the bytes are the ones the single-threaded loop made. The worker
is handed plain buffers in a DilateJob and makes nothing. tex_dilate_bytes is the slice-taking
form (safe_api.ludic), and examples/rendering/dilate.ludic holds the result against the old loop
(prints DILATE OK). It was 206 ms of the main thread in a Maroon Lake boot.