feat(render3d): compute and indirect draws on Vulkan, and tree layers culled on the GPU (opt-in)
- Device: multiDrawIndirect, drawIndirectFirstInstance and drawIndirectCount where present. - Buffers carry storage and indirect usage; a GPU-owned buffer is never swapped under a draw. - Compute programs from shaders/compute.list (binding 0 parameters, 1.. storage buffers), built by `ludic-dev shaders`; gpu_compute / gpu_dispatch / gpu_draw_mesh_indirect in gpu.ludic. - R3D_VK_PROBE=1: a dispatch read back (OK on the RTX 3070 Ti). - scatter_cull.comp: a tree layer's frustum test and LOD split on the GPU, with the lit, prepass, impostor and shadow-LOD draws reading its records. Behind R3D_GPU_CULL=1 and off by default: at the camp it is slower (43.0 fps against 53.3), because the frame's cost is per-draw descriptor sets and it adds empty-level draws. Validation-clean; OpenGL frames unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
04cda22d19
commit
c1eb5f399f
11 changed files with 463 additions and 15 deletions
10
packages/ludic.render3d/shaders/probe.comp
Normal file
10
packages/ludic.render3d/shaders/probe.comp
Normal file
|
|
@ -0,0 +1,10 @@
|
|||
// probe.comp - the compute path's own check (R3D_VK_PROBE=1): every float in the buffer times
|
||||
// the parameter, so a readback proves the dispatch, the bindings and the GPU-owned buffer.
|
||||
layout(local_size_x = 64) in;
|
||||
layout(set = 0, binding = 0) uniform Params { uint count; float mul; } pr;
|
||||
layout(set = 0, binding = 1) buffer Data { float v[]; } data;
|
||||
void main() {
|
||||
uint i = gl_GlobalInvocationID.x;
|
||||
if (i >= pr.count) { return; }
|
||||
data.v[i] = data.v[i] * pr.mul;
|
||||
}
|
||||
Loading…
Add table
Add a link
Reference in a new issue