Commit graph

2 commits

Author SHA1 Message Date
cc384efee0 feat(render3d): the meadow's blades are culled on the GPU - under 1.5 ms of a MoltenVK frame for all the grass
- grass_cull.comp decides each candidate blade once a frame (place, ground, density, water,
  slope, frustum, colour field) and writes survivors into three bands by distance, one indirect
  draw each; grass_inst.vert only bends and places the vertices. OpenGL keeps grass.vert's
  per-vertex path; R3D_GRASS_GPU=0 compares
- compute programs take sampled textures after their buffers (gpu_compute_tex / gpu_dispatch_tex),
  and every dispatch now records a compute-to-draw memory barrier
- the blades as a sward: 4 m cells, bands with five, three and one-quad blades, spacing doubling
  every 18 m to 70 m, never narrower than a pixel; lit facing the sun and leaning to the sky,
  shadow looked up above the ground (it read the terrain as its caster), a colour ramp that
  leaves only the sheath dark, clumps, dry patches and a tussock shade worked out per blade
- scatter layers flagged grass are skipped while the blades draw (and under R3D_NOGRASS);
  R3D_BLADES, R3D_NOBLADES win over the game's setting; R3D_GRASS_S0/D0 for measuring

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 15:47:36 +03:00
c1eb5f399f feat(render3d): compute and indirect draws on Vulkan, and tree layers culled on the GPU (opt-in)
- Device: multiDrawIndirect, drawIndirectFirstInstance and drawIndirectCount where present.
- Buffers carry storage and indirect usage; a GPU-owned buffer is never swapped under a draw.
- Compute programs from shaders/compute.list (binding 0 parameters, 1.. storage buffers),
  built by `ludic-dev shaders`; gpu_compute / gpu_dispatch / gpu_draw_mesh_indirect in gpu.ludic.
- R3D_VK_PROBE=1: a dispatch read back (OK on the RTX 3070 Ti).
- scatter_cull.comp: a tree layer's frustum test and LOD split on the GPU, with the lit, prepass,
  impostor and shadow-LOD draws reading its records. Behind R3D_GPU_CULL=1 and off by default:
  at the camp it is slower (43.0 fps against 53.3), because the frame's cost is per-draw
  descriptor sets and it adds empty-level draws. Validation-clean; OpenGL frames unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 14:19:43 +03:00