A chunk's dispatch was sized for its largest tile, so the far tiles beside a
near one ran thousands of empty invocations: 9.8 ms of grass at 4K on an RTX
3070 Ti. One dispatch per tile, sized to that tile, brings it to 5.0 ms - still
three times the chunked path's 1.7 (36 fps against 41), so GF_MESH_GRASS is not
implemented yet and the Advanced row says a coming update.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
NVIDIA Streamline's interposer exports none of VK_KHR_acceleration_structure's
commands, so ray tracing looks each up per device with vkGetDeviceProcAddr and
calls it through a pointer: create (four pointers -> result), destroy (device,
handle, allocator), build sizes (device, type, info, counts, sizes), the
command-buffer build (buffer, count, infos, ranges) and the device address
(device, info -> u64). Both runtimes assemble; nothing uses them yet.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The chunked grass path draws each chunk as one mesh-shader dispatch when the
setting asks and the card has VK_EXT_mesh_shader: work group y is a tile, x a
batch of 16 of its blades, and a blade the placement, density, frustum, water
or slope tests reject emits nothing - the instanced path still runs its eight
vertices to a degenerate position. grass.mesh generates the same blades as
grass.vert (grass_blade_mesh(4)'s rows, the same hashes, sway and lighting
normal), capped at the instanced path's 65535 a tile.
Vulkan: VK_EXT_mesh_shader with meshShader, and maintenance4 (glslang's mesh
stages declare LocalSizeId); vkCmdDrawMeshTasksEXT looked up per device, as the
Streamline interposer exports none; a *.mesh program's pipeline takes the mesh
stage and no vertex input, its bindings the mesh stage bit. gpu_has_mesh,
gpu_draw_mesh_tasks; r3d_mesh_grass and R3D_MESH_GRASS / R3D_NO_MESH.
bin/ludic-dev rebuilt: the committed binary predated the shader tool's mesh
support and compiled grass.mesh as a vertex stage.
PC (RTX 3070 Ti): the camp matches the chunked path; validation only the
no-window present-id message.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
HDR output: an HDR10 swapchain (A2B10G10R10, ST 2084 over BT.2020) when the
setting asks and the display offers it, with HDR metadata. The tonemap's HDR10
variant keeps the SDR picture up to a 200-nit paper white and rolls highlights
on to 1000 nits; the overlay's converts the interface to the same white. The
screen and LDR images go 10-bit with it; screenshots refuse while it is on.
OpenGL and the Vulkan SDR frame are unchanged. The instance asks for
VK_EXT_swapchain_colorspace. HDR metadata only where the loader has
vkSetHdrMetadataEXT: Streamline's interposer does not, and calling the thunk
crashed the game the moment the swapchain came up HDR10. PC 4K monitor: HDR10,
validation 0. R3D_HDR overrides the setting.
DLSS:
- the vertical jitter offset flips with Streamline's image (rows from the top):
unflipped, Quality resolved the ground into concentric rings;
- preset K in every mode: the default M put Performance at 18 ms a frame at 4K
on an RTX 3070 Ti (33 fps against 41 with DLSS off; with K, 60);
- the camera is jittered only while this frame holds a token and the last
evaluate worked.
R3D_DLSS_PRESET, R3D_CAM_LOG (the camera and DLSS state a frame) and
R3D_NOGRAIN for measuring.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ludic-dev shaders: a variant whose first file is *.mesh compiles that stage with
glslang -S mesh for Vulkan 1.3, without the vertex stage's depth-remap wrapper
or invariant gl_Position (a mesh shader writes an array of positions and
remaps depth itself). Its SPIR-V keeps the .vert name, so the manifest and the
loader are unchanged. The 51 existing programs build identical SPIR-V.
runtime: lsl_call_piii(fn, ptr, i32, i32, i32) calls a command-buffer command
through a pointer. NVIDIA Streamline's interposer exports no
vkCmdDrawMeshTasksEXT, so it is to be looked up per device with
vkGetDeviceProcAddr. Both runtimes assemble.
Nothing uses either yet.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
grass_flush uploaded u_tiles (vec4[256]) with u_fv(4n floats). On Vulkan an
array element is copied at the size given and placed at the array's stride,
so each tile got one float: nonsense corners and zero blades a cell. The
meadow had no grass on the Vulkan renderer since cdfffa6 while the draw
counts looked right. u_f4v uploads n vec4s; OpenGL was never affected.
PC camp: chunked path matches R3D_GRASS_TILES=1, validation 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
streamline.ludic: slInit before the Vulkan instance, feature support per
adapter, a frame token per frame, Reflex sleep and PCL latency markers, and
DLSS super resolution on the lit HDR frame (Halton jitter, depth + zero
motion vectors with camera motion from clipToPrevClip, matrices carrying
the Vulkan path's y flip and depth remap). Bloom, tonemap and sharpen read
the upscaled size. The device asks for privateData and present_id, which
Streamline's hooks need. R3D_DLSS / R3D_REFLEX / R3D_SL_LOG for tests.
gpu_feature_implemented: DLSS and Reflex.
ludic bundle (Windows): app native "<dir>" copies native libraries beside
the executable.
Verified on an RTX 3070 Ti: DLSS Quality evaluates 1280x720 -> 1920x1080.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
vk_sl_prefer(1) before Vk.open() makes the Windows loader try sl.interposer.dll
beside the executable and fall back to vulkan-1.dll. Thunks for the sl* exports
and for calling feature functions from slGetFeatureFunction; macOS stubs.
Verified on an RTX 3070 Ti: slInit eOk, DLSS, DLSS-RR, Reflex and PCL supported,
DLSS-G reports no supported adapter (needs RTX 40).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The water reflection keeps the main camera's frustum planes, and ac_visible tested the actor itself
against them: the town's people, far above the lake with their images far below the frame, all drew
again. The image (2L - y) is what is tested now - town, 69 skinned reflection draws -> 38, the frame
byte-identical; PC camp reflection scene CPU 793 -> about 760 us. R3D_PROF splits the reflection's
CPU into setup, terrain, scene and sky.
Tried and not kept, with the numbers (PC, camp, Vulkan GPU timestamps):
- the ground's cost is its scanned material taps: without them the terrain pass is 555 us of 1568,
without the photograph 1029; the sun, the noise fields and the grain are 20-100 us each. Moving the
cheap tier in from 200 m to 60 m changes nothing, so it is the far tier's single taps over the
mountains. Skipping the normal-map taps past 900 m (where the detail normal is fully faded) saved
22 us and moved a few pixels - reverted.
- grass: cutting its radius to 300 m barely moves it, so the cost is the near blades' pixels; moving
their three noise fields to the vertices changed nothing (747 us) - reverted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Where the frame goes, before optimising it further. R3D_PROF now prints the average frame: CPU before
the swap (the game's share apart), GPU and swap, and the renderer's CPU phases in frame order - the
shadow pass split into cascade fit, scatter casters and actor casters. On Vulkan the per-pass GPU
table comes from timestamp queries (host query reset asked for where the device has it); MoltenVK's
attribution is tile-based and not to be trusted per pass, the PC's is.
What it showed on the PC (camp): 4.3 ms CPU and 5.3 ms GPU a frame; the shadow pass was the largest
CPU phase (1.8 ms) and actors half of that. Every actor within 300 m was drawn into all five cascades,
though the outer two only shade receivers from 212 and 935 m out: 160 actors and 300 draws into each.
cast_band_reaches - the flowers' reach test, now shared - skips an actor for a cascade it cannot shade
(receivers counted from 0.85 of the previous split, where sunShadow's cross-fade begins).
PC camp, two runs each: mean frame 9553/9601 -> 8664/8656 us; CPU 4.3 -> 3.7 ms; shadow GPU 1.14 ->
0.79 ms; actor-shadow CPU 1.02 -> 0.60 ms; 2039 -> 1411 draws; self-tests 59/59, validation 0.
Mac: OpenGL shot viewpoints and the camp byte-identical; town 19 px at <= 2/255 on two flower stems a
few metres from the camera - the accepted leftover-binding difference, no shadow; self-tests 59/59 on
OpenGL and Vulkan. R3D_CAST_ALL=1 draws every caster into every cascade.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The reflection pass is a second copy of the terrain, the vegetation and the actors, and it ran on
every frame of a map with water - looking straight down at a meadow included. The mirrored body is cut
into rectangles where the ground lies below it (32 m over the height map in 8 x 8 blocks, the sea
beyond it coarse), and the pass runs when one meets the view frustum and its water is not all dry
where it shows, outside the frame or behind the ground. Three wrong turns on the way, each kept in a
comment: bounding spheres (a 156 m cell reached into the view from behind the camera), sight lines
that never asked whether the point was in the frame, and a walk of every rectangle every frame (0.1 s
per 400 frames on the PC until the blocks). Built once in 6.6 ms.
Mac, the five shot viewpoints: view c (the meadow) 225 reflection draws -> none; a, b, d, e unchanged;
OpenGL frames byte-identical; self-tests 59/59 on OpenGL and Vulkan; MoltenVK validation adds nothing.
PC camp (lake in view): 1882 draws either way, 3.8 s for 400 frames either way, three runs each,
self-tests 59/59, validation 0. R3D_REFL_ALWAYS=1 runs the pass every frame; R3D_REFL_DBG=1 prints the
test once.
R3D_ACTOR_CENSUS=<frame> prints the lit pass's actors grouped by model: town is 17 models and 69 draws,
and only four rigid models repeat (17 actors) - too little for instancing to be worth a shader variant.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
tersun.frag rasterised every terrain patch a second time into a screen buffer, because on OpenGL the
cascade read inside terrain.frag fell off a driver cliff (about 6 ms a frame). Vulkan has no such
cliff: its terrain programs are the SUN_INLINE variant, which evaluates the same two tiers with the
same normal and cross-fade in the ground's own shader, and terrain_draw skips the pass - in the frame
and in the water reflection. OpenGL keeps the pass. R3D_SUN_PASS=1 keeps it on Vulkan, for comparing.
Mac Vulkan town 1759 -> 1692 draws; frames within 2/255 of the pass (15331 px, float rounding). PC camp
1984 -> 1882 draws, 3.9 s for 400 frames either way, self-tests 59/59, validation 0. OpenGL frames
byte-identical.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A panel and its label alternated the white and font textures, and every switch closed the overlay's
draw range: 77-91 ranges a frame in town. The font atlas is now bound beside the image texture for the
whole flush, and v carries 4 x the vertex's mode (0 image, 1 font, 2 flat colour) on top of its real
coordinate, so only a different image texture or a clip rectangle closes a range. Town: 24 ranges.
OpenGL frames byte-identical at all five viewpoints; MoltenVK with validation adds no message.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Grass (Vulkan with multi-draw indirect): every visible tile is a record in one buffer, uploaded once a
frame, and each band draws its records 256 at a time. A record's firstInstance is its place in the
chunk times 65536; grass.vert's TILES variant reads that place's corner and indices per cell from
u_tiles. R3D_GRASS_TILES=1 keeps a draw per tile. OpenGL is unchanged.
Casters: a LOD level with no impostor is drawn into a shadow cascade only when its distance band,
widened by six times its height, the camera's height over the ground and the frustum's corner reach,
can touch that cascade's receivers. The flowers' mesh levels (6 - 30 m) leave the three outer
cascades. R3D_CAST_ALL=1 draws every level everywhere. OpenGL frames byte-identical at all five
viewpoints; alpha-tested shadow draws at a 460 -> 244.
gpu_has_mdi() guards both this and the GPU-culled trees' multi-record draws.
Camp: Mac Vulkan 2645 -> 2191 (grass) -> 2034 draws; PC 2657 -> 2046, self-tests 59/59, validation 0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every level of a kit tree or rock carries the same materials, so a GPU-culled layer merges each
material's levels into one mesh (layer_arena_build) and scatter_cull.comp writes a record per
(material, level) naming that level's index and vertex range; one indirect draw covers them all.
A conifer's lit pass and prepass go from 8 draws to 2, its shadow LOD from 6 to 2. Layers that do
not fit (card levels, other materials or attributes, 32-bit indices) keep the CPU path.
GPU culling is now the Vulkan default (R3D_GPU_CULL=0 turns it off). Camp benchmark, 400 frames:
PC 2791 -> 2657 draws, 4.3 -> 4.1 s (both runs); Mac 2791 -> 2657, 8.0/7.4 -> 7.7/7.2 s.
Self-tests 59/59 on both machines, PC validation 0 errors; frames within run-to-run noise; OpenGL
frames unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
gpu_rb_storage makes the image with the samples asked for; a pass takes its attachments' count and
every pipeline in it matches; a blit from a multisampled source resolves on the end of an empty
dynamic-rendering pass (colour averaged, depth from sample zero - vkCmdResolveImage cannot do depth).
gpu_msaa_max reads the device's colour-and-depth sample limits, and post_set_msaa clamps to it, so
the Anti-aliasing setting is live on Vulkan. The pipeline cache's pass key no longer packs samples
into three bits.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
gpu_feature_implemented(GF_VULKAN) is true: the Vulkan renderer draws the whole game (validation
clean, 105 fps at the camp on the RTX 3070 Ti). Ray tracing, DLSS, Reflex, HDR output and mesh-shader
grass still report false, so their rows keep saying they take effect later.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- A block sub-allocator: 64 MB blocks per memory type, images and buffers kept apart, first fit with
alignment, freed ranges merged and empty blocks given back; anything over 16 MB still gets its own
allocation. The self-tests' second world holds 94 allocations instead of 8735 (the driver's limit
refused a shadow map before). Camp bench unchanged, 105.3 fps.
- shadow_set_res keeps the size that worked when the card has no memory for the new one, and records
it in shadow_refused, instead of ending with no shadow map; gpu_tex_ok says whether a texture has an
image behind it.
- Image barriers skip an image that was never made (a failed allocation used to crash there), and
R3D_VK_ERRLOG=<file> appends every Vulkan failure line by line, so a crash no longer takes the
message with it.
- R3D_VK_PROF reports draws asked for and not made, so a layer missing from a frame is never silent.
Validation on (VK_INSTANCE_LAYERS): the game's self-tests 61 OK, 0 errors, on the RTX 3070 Ti.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- A draw's set carries its textures and lives in a pool of its own, keyed by each texture's handle,
generation and sampler; the uniform blocks are dynamic uniform buffers into the frame's ring,
bound with the draw's offsets, so a draw that changes only uniforms allocates and writes no set.
The sampler for an unbound slot is made once instead of looked up by string every draw.
Camp bench, RTX 3070 Ti, 400 frames, twice: 102.6 fps (53.3 before; OpenGL 114.3), sets 0.9 ms
against 8.5. MoltenVK (M4 Pro): sets 0.2 ms.
- Integer vertex attributes read by float inputs use USCALED formats (a_joints was UINT against a
vec4), and every shader input a mesh does not feed reads a shared zero buffer. NVIDIA drew
anyway; MoltenVK refused every skinned pipeline and the whole actor layer was missing from the
Mac's Vulkan frame. A draw skipped for want of a pipeline now says so, once per program.
- A failed image allocation is reported instead of bound as a null allocation.
Validation proven on (VK_INSTANCE_LAYERS): 0 errors over the game's self-tests (61 OK) and a frame.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- sky_start_yaw: a game that turns the sky at boot sets it before r3d_init and the image-based light
is baked once at that yaw; sky_set_yaw to a yaw already baked bakes nothing.
- terrain_init_step: the height field and normals, the sun shadow, the materials, the patches and
programs as separate steps; r3d_load_count is 8, so a loading bar moves through the terrain (half
the start-up) instead of jumping over it. terrain_init runs the four in a row as before.
- The per-cascade actor cull centres its sphere at half the model's height and sizes it by the
larger of height and radius: centred at the radius, it dropped a head at a cascade's edge (11 px
in town, found by the drawstats comparison). Lake was byte-identical before and after.
OpenGL frames at the five viewpoints unchanged; the game's 59 self-tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
actor_draw_casters drew every visible actor within 300 m into all five cascades. The draw-call
count (R3D_DRAWSTATS) found about 1500 skinned shadow draws a frame in town against 69 in the lit
pass. The cascade is an orthographic box, so a bounding sphere (1.5x the actor's radius) outside
its clip x/y casts nothing into that layer and is skipped. OpenGL frames at the five viewpoints
are unchanged; the game's 59 self-tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- r3d_open opens the window with nothing baked; r3d_load_count / r3d_load_step run the rest (sky
and light, terrain, shadow and screen targets, cover and actors, grass) so a game can present a
loading frame between them. r3d_init is the same calls in a row.
- sky_set_quality(width): the prefiltered sky light at 256 / 512 / 1024, baked again.
- post_set_msaa(samples): multisampling on OpenGL, remade in place (post_msaa_live is false on
Vulkan); post_free frees the multisampled framebuffer too.
- STREAM_BUDGET_US is a variable; r3d_fog_scale multiplies the fog the day sets.
- Vulkan buffers carry UNIFORM_BUFFER usage: the frame's ring is bound as uniform buffers and was
created without it (VUID-VkWriteDescriptorSet-descriptorType-00330, found on MoltenVK).
OpenGL frames at the five viewpoints unchanged; the game's 59 self-tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Measurement only: a tally at each of gpu.ludic's five draw doors, program
changes and texture binds, keyed by the open profiler pass. Off unless
R3D_DRAWSTATS is set; frames byte-identical with it off and on (Mac, OpenGL).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- r3d_set_anisotropy(level) updates every mipmapped texture already loaded (OpenGL parameter,
Vulkan sampler record), not only later uploads; the game's setting never reached the scanned
materials, which load before the settings are read.
- shadow_set_res(size) remakes the cascades at 1024 / 2048 / 4096 (shadow_res replaces the
SHADOW_RES constant); lighting.glsl reads the texel size from the map, so OpenGL frames at
2048 are unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- Device: multiDrawIndirect, drawIndirectFirstInstance and drawIndirectCount where present.
- Buffers carry storage and indirect usage; a GPU-owned buffer is never swapped under a draw.
- Compute programs from shaders/compute.list (binding 0 parameters, 1.. storage buffers),
built by `ludic-dev shaders`; gpu_compute / gpu_dispatch / gpu_draw_mesh_indirect in gpu.ludic.
- R3D_VK_PROBE=1: a dispatch read back (OK on the RTX 3070 Ti).
- scatter_cull.comp: a tree layer's frustum test and LOD split on the GPU, with the lit, prepass,
impostor and shadow-LOD draws reading its records. Behind R3D_GPU_CULL=1 and off by default:
at the camp it is slower (43.0 fps against 53.3), because the frame's cost is per-draw
descriptor sets and it adds empty-level draws. Validation-clean; OpenGL frames unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`fn name` read `while p < fn and ...` as a reference to a function called `and`, which broke
tools/ludic-cli/glgen.ludic and so the dev tool's own build. A reference now needs the name on
the same line and not one of and / or / not / in / is. The threads example checks `fn` as a
local before `and`. Reseeded; bootstrap-cfree reproduces the seed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- `fn name` names a top-level function as a value (E_FNREF, lowers to @fn_<name>); the worker
entry point for Job.parallel_for, which checks it takes (int, pointer-like) and returns void.
- runtime/native/threads.ll (pthreads) and threads_win.ll (Win32 SRWLOCK/CONDITION_VARIABLE): a
pool of one worker per core but one, parked between batches; every thread claims chunks by
compare-and-swap. Linked only into programs that use Job/Promise/Sync, by `ludicc -o`,
`ludic build` and the test suite's build helper.
- Sync.* is real: native mutexes, atomics as cmpxchg retry loops (neither clang takes atomicrw,
the PC's rejects seq_consistent), mutex-guarded channels, Sync.cpu_count from the OS.
- spawn/despawn on a pool thread stop the program with a located panic.
- examples/library/threads.ludic and its test; docs for fn, Job.parallel_for, Job.is_worker.
- Reseeded (bootstrap-cfree: out.ll == seed.ll). 141/141 on macOS; jobs, threads and the guard
pass on Windows from the reseeded Windows seed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R3D_VK_PROF at the camp view (about 2950 draws a frame) showed the frame spent in bookkeeping:
19.5 ms a frame finding pipelines by string key and 16 ms building descriptor sets.
- Pipelines: a mesh carries an interned vertex-layout id (dropped only when an attribute's shape
changes, not when an instance buffer is swapped), render state packs into an int and the pass
formats into another; the program's last hit is tried first. The string path only builds.
- Samplers: each texture keeps the sampler for its parameters until they change.
- Descriptor sets: a program's last set is reused within the frame while its blocks and resolved
textures are unchanged; a uniform write that repeats the value it already holds changes nothing.
Headless on the RTX 3070 Ti at 1920x1080: 21.7 -> 53.3 fps (pipelines 0.2 ms, sets 8.5 ms, inside
draws 10 ms a frame; OpenGL 114 fps). The camp frame is unchanged and validation-clean. OpenGL frames
byte-identical at the five viewpoints; 59 self-tests pass; VKRES OK.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- A buffer re-uploaded after a draw this frame read it gets fresh storage, as OpenGL orphans
one; storage moved off or freed while the frame still reads it is destroyed after the frame's
submit. Draws mark the buffers they bind. No flush for buffer work.
- Mipmaps asked for mid-frame (the exposure measure, every frame) are recorded into the open frame
after its pass; growing a chain the first time keeps its one-shot path.
- R3D_VK_PROF prints, every 120 frames, draws and flushes per frame and the milliseconds spent
finding pipelines, filling descriptor sets and inside draws.
The gain was small - the camp view headless at 1920x1080 on the RTX 3070 Ti went from 21.3 to
21.7 fps (OpenGL: 114 fps) - so these flushes were not what holds the frame; the profile is how
the rest is found. The frame is unchanged and validation-clean; OpenGL frames byte-identical at the
five viewpoints with 59 self-tests passing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- gpu_select takes Vulkan for a window on Windows; gvk_init asks for VK_KHR_surface,
VK_KHR_win32_surface and VK_KHR_swapchain when the renderer will present. A window elsewhere
stays on OpenGL with the reason.
- gvk_open opens the window without a GL context and makes the screen images at its client area;
gvk_swap_make builds the swapchain on the Win32 surface (FIFO with vsync, mailbox or immediate
without) and rebuilds it when it is out of date, suboptimal or the window changes size.
- gvk_present blits the screen image into the acquired swapchain image, flipped (the screen keeps
OpenGL's bottom-up rows), and presents it.
- Samplers pointed at a texture unit with u_i and then fed by binding units - the actors do this -
read that unit's texture; on Vulkan they drew with the white stand-in (the tent, the log, the
chair, the workbench).
- gl.ll carries a weak win_gl_drawable for headless macOS builds.
On the RTX 3070 Ti the windowed build runs the valley through Vulkan at 3840x2160 with the HUD and
the frame-rate readout (16 fps: every upload still waits on the frame, and buffers are
host-visible). The headless Vulkan frame is unchanged; OpenGL frames byte-identical at the five
viewpoints with 59 self-tests passing. The actor fix is compiled and OpenGL-verified, not yet
seen on the PC; blades at the grass's middle distance still draw black on Vulkan.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
OpenGL links a vertex output to a fragment input by name; SPIR-V links them by location, and
glslang's --auto-map-locations numbered each stage in its own declaration order. The foliage
prepass's depth.frag declares v_wpos then v_uv where model.vert writes v_wpos, v_nrm, v_uv, so
the prepass read a normal as its texture coordinate, its alpha test cut every flower head and
leaf, and the lit pass (depth EQUAL) drew nothing over them. `ludic-dev shaders` now gives both
stages explicit locations: the vertex stage's out order numbers them and the fragment stage looks
each in up by name. All 45 variants checked: every fragment input sits on its vertex output.
Also:
- A clear still waiting for its pass when the framebuffer changes now runs on that framebuffer,
instead of becoming the load op of whichever pass began next.
- R3D_DUMP_ATLAS writes every impostor and card atlas a run bakes (build/atlas_<n>_*.ppm); the
40 baked on Vulkan match OpenGL's.
The PC's Vulkan frame now shows the flowers as OpenGL does, validation-clean. ludic-dev test 140
passed; OpenGL frames byte-identical at the five viewpoints; 59 self-tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three things the first Vulkan frames on the RTX 3070 Ti showed against OpenGL on the same PC:
- Exposure: the adaptation pass reads the HDR scene's smallest mip, and a render target made
without pixels had one level, so exposure came from a single texel. A target asked for mipmaps
now grows a full chain (level 0 kept), and passes draw through level-0 views keyed by the
image's generation.
- Foliage: the depth prepass and the lit pass (depth EQUAL) are different variants. Vulkan vertex
stages now declare an invariant gl_Position so both land on the same depth.
- Alpha to coverage is enabled only on a multisampled pass. OpenGL ignores it without MSAA; Vulkan
with one sample dropped every fragment under half alpha.
vk_resources also checks a big-endian 16-bit RGB upload (a normal map). VKRES OK; ludic-dev test
140 passed; the PC's Vulkan frame is validation-clean; OpenGL frames byte-identical at the five
viewpoints with 59 self-tests passing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
gpu.ludic's calls branch to the Vulkan backend when it was chosen and came up; every OpenGL
statement is unchanged, only guarded. R3D_GFX=vk selects it in a headless run (the window's
swapchain is the next milestone) and falls back to OpenGL, with the reason, when the device or
the SPIR-V manifest is missing.
- Render state is cached as before and turned into pipelines at the draw; u_* and sampler binds
go into the variant's uniform blocks; meshes, buffers and textures are gvk_* objects;
framebuffer binds are dynamic-rendering passes; same-size blits are image copies; the screen,
the photograph read-back, the present and the screenshot go through the frame.
- Work that submits on its own (uploads, read-backs, new or freed images and buffers) flushes the
frame first, so it runs in OpenGL's order. A read may take fewer channels than the image has
(the height field's R from its RGBA32F bake). Pipeline keys name vertex bindings by order, not
buffer handle, so re-pointed instance buffers keep their pipeline.
The valley renders at frame 90 validation-clean on the RTX 3070 Ti and on MoltenVK. OpenGL frames
byte-identical at the five viewpoints; 59 self-tests pass with no GL error.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The drawing half of the Vulkan backend, compiled into render3d and not yet reached from it:
- gvk_program: a program handle as its manifest variant - two SPIR-V modules, a descriptor set
layout (the vertex block at 0, the fragment block at 1, the samplers at their manifest bindings)
and a pipeline layout. Nothing is compiled at run time.
- gvk_pipeline: one pipeline per program, recorded vertex layout, render state and pass formats,
built the first time that combination draws. Attributes the shader does not read are left out;
the front face is clockwise, since neither API flips y between clip space and its target rows.
- gvk_uniform / gvk_u_set / gvk_bind_texture: loose uniforms written into each stage's block at
the manifest's offsets and array strides; gvk_draw_set copies the blocks into a per-frame ring
at the device's alignment and fills a descriptor set from a per-frame pool, with a white 1x1
texture for a sampler nothing was bound to.
- The frame: one command buffer; a framebuffer bind ends the pass and the next begins at its first
clear or draw (a clear that comes first is the load op); attachments move to attachment layouts
for the pass and back to SHADER_READ_ONLY after it, a cascade drawn through a view of its layer.
gvk_present submits and waits; gvk_screenshot reads the screen image back bottom row first.
r3d.ludic now imports gpu_manifest.ludic too. Textures remember their size.
OpenGL frames byte-identical at the five viewpoints; 59 self-tests pass; VKRES OK and VKDEVICE OK.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The renderer's projections are OpenGL's, whose clip-space depth runs from -w to w; Vulkan clips
everything below 0. `ludic-dev shaders` now wraps each vertex stage - its own main runs, then
gl_Position.z = (z + w) / 2 - so every variant's depth lands in [0, w]. The GLSL the OpenGL
renderer compiles is untouched. No y flip is needed: a Vulkan target's row 0 is where OpenGL's is
(NDC y = -1), so render to texture, sampling and gl_FragCoord agree between the two, and only the
present and the screenshot flip.
45 vertex modules regenerated (spirv-val clean, manifest unchanged); ludic-dev test 140 passed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
gpu_vk_res.ludic gains what gpu.ludic's texture and buffer handles stand for on Vulkan:
- gvk_tex_storage / _upload / _mips / _read / _release: an image and view per handle, a full mip
chain when pixels come with it (the renderer asks for mipmaps after the upload) and one level
for a target, every level in SHADER_READ_ONLY between uses. Uploads are converted to what the
image stores: a missing alpha filled opaque (three-channel formats are stored with four), 16-bit
PNG samples byte-swapped when the unpack state says so, 32-bit float HDR halved into half
floats. Mips are blitted down level by level; a read-back brings level 0 home.
- gvk_sampler: one VkSampler per filter / wrap / compare / anisotropy combination, made when first
asked for, with GL's defaults where the renderer set nothing.
- gvk_buf_*: vertex, index and instance buffers behind one handle, kept when an upload fits.
Host-visible while the backend comes up.
gpu_vk.ludic switches on anisotropic sampling where the device has it and reads its limit.
examples/rendering/vk_resources.ludic checks it all: VKRES OK, validation-clean, every allocation
freed, on the RTX 3070 Ti (16x anisotropy) and on MoltenVK. One run on the Mac crashed while a
headless game run was using the GPU and did not come back in two reruns. OpenGL frames
byte-identical at the five viewpoints; 59 self-tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The Vulkan backend is two files. gpu_vk.ludic is the device, memory and one-shot commands,
pure Vulkan, and still runs on its own (vk_device.ludic). gpu_vk_res.ludic is the resources in
the renderer's vocabulary - OpenGL's names for formats, filters and blend factors, which only
exist where Gl.* is named - starting with the format table. r3d.ludic imports both after
gpu.ludic; nothing calls them yet.
OpenGL frames byte-identical at the five viewpoints; 59 self-tests pass; VKDEVICE OK.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first part of the Vulkan side of gpu.ludic, not yet reached from the renderer: gvk_init
brings up the instance (portability enumeration where offered), the first discrete GPU, a
graphics queue and a device with the Tier 1 floor switched on (dynamic rendering,
synchronization2, descriptor indexing, timeline semaphores), and says why when it cannot so
the caller stays on OpenGL. gvk_mem_type / gvk_alloc pick and allocate memory (one allocation
per resource while the backend comes up), and gvk_once_begin / gvk_once_end carry uploads,
bakes and read-backs through submit and wait.
examples/rendering/vk_device.ludic runs it alone: VKDEVICE OK and validation-clean on the RTX
3070 Ti and on MoltenVK.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A game ended its frame with Gl.swap() and shot it with Gl.screenshot, which tied it to OpenGL
and made it name Gl.* only so the parser would splice the GL runtime. gpu_present and
gpu_screenshot carry both (and name Gl.* themselves, so importing render3d is enough), and
r3d_present / r3d_screenshot are what a game calls. The Vulkan backend takes both over.
OpenGL frames byte-identical at the five viewpoints; 59 self-tests pass with no GL error.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A Slang vertex and fragment shader draw a triangle into an image with dynamic rendering (no
render pass object: the pipeline names its colour format), image barriers move it to the
attachment and transfer layouts, and vkCmdCopyImageToBuffer reads it back into a PPM. It is
the path the Vulkan renderer's passes and screenshots take.
The corner comes from SV_VulkanVertexID: Slang compiles SV_VertexID to gl_VertexIndex -
gl_BaseVertex, which declares the draw-parameters capability.
Validation-clean and VKTRIANGLE OK on the RTX 3070 Ti (Windows) and the M4 Pro (MoltenVK),
13824 pixels covered on both.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The last of milestone 1: no OpenGL call is left in render3d outside gpu.ludic but the clock
and the CPU-side buffer helpers.
- gpu_program builds a program and remembers the variant it came from (vertex, fragment and
defines as one line - the SPIR-V manifest's key), so a backend that cannot compile at run
time finds the pipeline for the same handle. gpu_use_program and gpu_program_free replace
32 uses and 2 frees; the overlay's program is recorded under overlay.vert|overlay.frag.
- gpu_query_new / _begin / _end / _result carry R3D_PROF's timers.
- gpu_open, gpu_vsync, gpu_renderer_name and gpu_resize_check carry the context.
- r3d_program_tess is gone: nothing called it and no tessellation shader exists.
- R3D_GLCHECK also checks after framebuffer and renderbuffer changes, viewport, draw buffers,
program use, texture parameters, the resize check and between frames.
OpenGL frames are byte-identical at the five viewpoints; 59 self-tests pass with and without
R3D_GLCHECK, with no GL error; ludic-dev test 140 passed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The overlay flushed into whatever framebuffer was bound last, which raised GL error 1286 on
every run: ov_flush now binds the screen and its viewport before it draws.
R3D_GLCHECK=1 checks each draw, clear, blit, upload and attachment for a pending error or an
incomplete framebuffer and names the target, and each sampler bind for a texture with no image
or a mipmap filter without mipmaps. Off, it costs one flag test.
Still open: an intermittent 1286 reported at "terrain shadow bake" after the self-tests (about
half the runs), and one macOS "unloadable" texture warning at shutdown while the post targets are
freed. No frame samples a bad texture.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
win_native_window / win_native_instance hand a Vulkan swapchain what it is created on: the
HWND and module instance on Windows (VkWin32SurfaceCreateInfoKHR), the game's view on macOS.
Headless builds define both as null (weak on macOS, so a windowed link keeps cocoa.ll's), so
code that asks for them links everywhere and simply finds no window.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
gpu_manifest.ludic loads shaders/spv/manifest.txt and answers what a backend asks while drawing:
which variant a program built by r3d_program is, where a uniform lives in its stage's block,
which binding a sampler has, what the vertex inputs are. examples/rendering/vk_manifest.ludic
resolves every program in variants.list against it (45 of 45) and checks a few offsets and
bindings. Nothing imports it yet.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Framebuffers and their attachments, renderbuffers (multisampled too), draw and read buffers,
completeness checks, blits, viewports, clears, the screen framebuffer, the multisample enable,
the wireframe switch and GL error checks now go through gpu_fb_* / gpu_rb_* / gpu_viewport /
gpu_clear / gpu_blit / gpu_check, and no other file names them. Each call is the one GL call it
replaces, in the same order: the fixed viewpoints render bit-identically and the game's
self-tests report exactly what they did.
What is attached to each framebuffer - colour slots, a depth texture or one layer of an array,
renderbuffers and their samples - is recorded as it is attached, for a backend that builds
render passes and image views. gpu_read_screen is the frame read-back a photograph takes.
R3D_GLCHECK=1 checks, around every draw, clear and blit, that the bound framebuffer is complete
and that no error is left behind, naming the framebuffer. It has already narrowed the old
"gl error 1286 at terrain shadow bake": the error is pending before a draw after the map self-
test, so it comes from a call that is not a draw, clear or blit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Texture creation, 2D and array uploads, pixel packing, filters, wraps, comparison, border,
anisotropy, mipmaps, texture units, read-backs and freeing now go through gpu_tex_* and
gpu_bind_sampler, and no other file names them. Each call is the one GL call it replaces,
in the same order, so OpenGL renders bit-identically at the fixed viewpoints and the game's
self-tests report what they did before.
What those calls say about a texture - size, format, layers, min and mag filter, wraps,
comparison, mipmaps, anisotropy - is recorded per handle as it is set: a backend with
immutable images and separate sampler objects creates both from exactly that.
r3d_bind_tex takes a texture kind (GPU_TEX2D / GPU_TEX2D_ARRAY) instead of a GL target.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every vertex array, vertex and index buffer, attribute pointer, instance divisor, stream
upload and draw call now goes through gpu_mesh_* / gpu_buffer_* / gpu_draw_*, and no other
file in the package names them. A Mesh records its layout as it is built - which buffer
feeds which attribute at what stride and offset, per vertex or per instance - so a backend
that bakes vertex input into a pipeline can read it back. On OpenGL each call is the GL it
replaces, in the same order: the five fixed viewpoints render bit-identically and the game's
self-tests report exactly what they did before.
scatter_attach takes the mesh rather than its vertex array; a mesh now frees every vertex
buffer it owns (glTF meshes used to keep all but the first); the two helpers nothing called,
mesh_grid_patches and mesh_instance_buffer, are gone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>