Every level of a kit tree or rock carries the same materials, so a GPU-culled layer merges each
material's levels into one mesh (layer_arena_build) and scatter_cull.comp writes a record per
(material, level) naming that level's index and vertex range; one indirect draw covers them all.
A conifer's lit pass and prepass go from 8 draws to 2, its shadow LOD from 6 to 2. Layers that do
not fit (card levels, other materials or attributes, 32-bit indices) keep the CPU path.
GPU culling is now the Vulkan default (R3D_GPU_CULL=0 turns it off). Camp benchmark, 400 frames:
PC 2791 -> 2657 draws, 4.3 -> 4.1 s (both runs); Mac 2791 -> 2657, 8.0/7.4 -> 7.7/7.2 s.
Self-tests 59/59 on both machines, PC validation 0 errors; frames within run-to-run noise; OpenGL
frames unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- Device: multiDrawIndirect, drawIndirectFirstInstance and drawIndirectCount where present.
- Buffers carry storage and indirect usage; a GPU-owned buffer is never swapped under a draw.
- Compute programs from shaders/compute.list (binding 0 parameters, 1.. storage buffers),
built by `ludic-dev shaders`; gpu_compute / gpu_dispatch / gpu_draw_mesh_indirect in gpu.ludic.
- R3D_VK_PROBE=1: a dispatch read back (OK on the RTX 3070 Ti).
- scatter_cull.comp: a tree layer's frustum test and LOD split on the GPU, with the lit, prepass,
impostor and shadow-LOD draws reading its records. Behind R3D_GPU_CULL=1 and off by default:
at the camp it is slower (43.0 fps against 53.3), because the frame's cost is per-draw
descriptor sets and it adds empty-level draws. Validation-clean; OpenGL frames unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>