- Device: multiDrawIndirect, drawIndirectFirstInstance and drawIndirectCount where present. - Buffers carry storage and indirect usage; a GPU-owned buffer is never swapped under a draw. - Compute programs from shaders/compute.list (binding 0 parameters, 1.. storage buffers), built by `ludic-dev shaders`; gpu_compute / gpu_dispatch / gpu_draw_mesh_indirect in gpu.ludic. - R3D_VK_PROBE=1: a dispatch read back (OK on the RTX 3070 Ti). - scatter_cull.comp: a tree layer's frustum test and LOD split on the GPU, with the lit, prepass, impostor and shadow-LOD draws reading its records. Behind R3D_GPU_CULL=1 and off by default: at the camp it is slower (43.0 fps against 53.3), because the frame's cost is per-draw descriptor sets and it adds empty-level draws. Validation-clean; OpenGL frames unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
10 lines
473 B
Text
10 lines
473 B
Text
// probe.comp - the compute path's own check (R3D_VK_PROBE=1): every float in the buffer times
|
|
// the parameter, so a readback proves the dispatch, the bindings and the GPU-owned buffer.
|
|
layout(local_size_x = 64) in;
|
|
layout(set = 0, binding = 0) uniform Params { uint count; float mul; } pr;
|
|
layout(set = 0, binding = 1) buffer Data { float v[]; } data;
|
|
void main() {
|
|
uint i = gl_GlobalInvocationID.x;
|
|
if (i >= pr.count) { return; }
|
|
data.v[i] = data.v[i] * pr.mul;
|
|
}
|