r3d_hdr_calibrate(peak, paper, black) feeds the tonemap's and the overlay's HDR10 variants and
the display's HDR metadata; ov_hdr_nits draws a calibration patch at a number of nits.
gpu_caps_probe asks the running Vulkan renderer's instance instead of making and destroying a
second one under Streamline's interposer, which left the next swapchain rebuild calling address 0.
The overlay gets an HDR10 variant, and the tonemap brightens HDR highlights by one factor rather
than per channel.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The chunked grass path draws each chunk as one mesh-shader dispatch when the
setting asks and the card has VK_EXT_mesh_shader: work group y is a tile, x a
batch of 16 of its blades, and a blade the placement, density, frustum, water
or slope tests reject emits nothing - the instanced path still runs its eight
vertices to a degenerate position. grass.mesh generates the same blades as
grass.vert (grass_blade_mesh(4)'s rows, the same hashes, sway and lighting
normal), capped at the instanced path's 65535 a tile.
Vulkan: VK_EXT_mesh_shader with meshShader, and maintenance4 (glslang's mesh
stages declare LocalSizeId); vkCmdDrawMeshTasksEXT looked up per device, as the
Streamline interposer exports none; a *.mesh program's pipeline takes the mesh
stage and no vertex input, its bindings the mesh stage bit. gpu_has_mesh,
gpu_draw_mesh_tasks; r3d_mesh_grass and R3D_MESH_GRASS / R3D_NO_MESH.
bin/ludic-dev rebuilt: the committed binary predated the shader tool's mesh
support and compiled grass.mesh as a vertex stage.
PC (RTX 3070 Ti): the camp matches the chunked path; validation only the
no-window present-id message.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
HDR output: an HDR10 swapchain (A2B10G10R10, ST 2084 over BT.2020) when the
setting asks and the display offers it, with HDR metadata. The tonemap's HDR10
variant keeps the SDR picture up to a 200-nit paper white and rolls highlights
on to 1000 nits; the overlay's converts the interface to the same white. The
screen and LDR images go 10-bit with it; screenshots refuse while it is on.
OpenGL and the Vulkan SDR frame are unchanged. The instance asks for
VK_EXT_swapchain_colorspace. HDR metadata only where the loader has
vkSetHdrMetadataEXT: Streamline's interposer does not, and calling the thunk
crashed the game the moment the swapchain came up HDR10. PC 4K monitor: HDR10,
validation 0. R3D_HDR overrides the setting.
DLSS:
- the vertical jitter offset flips with Streamline's image (rows from the top):
unflipped, Quality resolved the ground into concentric rings;
- preset K in every mode: the default M put Performance at 18 ms a frame at 4K
on an RTX 3070 Ti (33 fps against 41 with DLSS off; with K, 60);
- the camera is jittered only while this frame holds a token and the last
evaluate worked.
R3D_DLSS_PRESET, R3D_CAM_LOG (the camera and DLSS state a frame) and
R3D_NOGRAIN for measuring.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
streamline.ludic: slInit before the Vulkan instance, feature support per
adapter, a frame token per frame, Reflex sleep and PCL latency markers, and
DLSS super resolution on the lit HDR frame (Halton jitter, depth + zero
motion vectors with camera motion from clipToPrevClip, matrices carrying
the Vulkan path's y flip and depth remap). Bloom, tonemap and sharpen read
the upscaled size. The device asks for privateData and present_id, which
Streamline's hooks need. R3D_DLSS / R3D_REFLEX / R3D_SL_LOG for tests.
gpu_feature_implemented: DLSS and Reflex.
ludic bundle (Windows): app native "<dir>" copies native libraries beside
the executable.
Verified on an RTX 3070 Ti: DLSS Quality evaluates 1280x720 -> 1920x1080.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
gpu_rb_storage makes the image with the samples asked for; a pass takes its attachments' count and
every pipeline in it matches; a blit from a multisampled source resolves on the end of an empty
dynamic-rendering pass (colour averaged, depth from sample zero - vkCmdResolveImage cannot do depth).
gpu_msaa_max reads the device's colour-and-depth sample limits, and post_set_msaa clamps to it, so
the Anti-aliasing setting is live on Vulkan. The pipeline cache's pass key no longer packs samples
into three bits.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- A block sub-allocator: 64 MB blocks per memory type, images and buffers kept apart, first fit with
alignment, freed ranges merged and empty blocks given back; anything over 16 MB still gets its own
allocation. The self-tests' second world holds 94 allocations instead of 8735 (the driver's limit
refused a shadow map before). Camp bench unchanged, 105.3 fps.
- shadow_set_res keeps the size that worked when the card has no memory for the new one, and records
it in shadow_refused, instead of ending with no shadow map; gpu_tex_ok says whether a texture has an
image behind it.
- Image barriers skip an image that was never made (a failed allocation used to crash there), and
R3D_VK_ERRLOG=<file> appends every Vulkan failure line by line, so a crash no longer takes the
message with it.
- R3D_VK_PROF reports draws asked for and not made, so a layer missing from a frame is never silent.
Validation on (VK_INSTANCE_LAYERS): the game's self-tests 61 OK, 0 errors, on the RTX 3070 Ti.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- A draw's set carries its textures and lives in a pool of its own, keyed by each texture's handle,
generation and sampler; the uniform blocks are dynamic uniform buffers into the frame's ring,
bound with the draw's offsets, so a draw that changes only uniforms allocates and writes no set.
The sampler for an unbound slot is made once instead of looked up by string every draw.
Camp bench, RTX 3070 Ti, 400 frames, twice: 102.6 fps (53.3 before; OpenGL 114.3), sets 0.9 ms
against 8.5. MoltenVK (M4 Pro): sets 0.2 ms.
- Integer vertex attributes read by float inputs use USCALED formats (a_joints was UINT against a
vec4), and every shader input a mesh does not feed reads a shared zero buffer. NVIDIA drew
anyway; MoltenVK refused every skinned pipeline and the whole actor layer was missing from the
Mac's Vulkan frame. A draw skipped for want of a pipeline now says so, once per program.
- A failed image allocation is reported instead of bound as a null allocation.
Validation proven on (VK_INSTANCE_LAYERS): 0 errors over the game's self-tests (61 OK) and a frame.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- Device: multiDrawIndirect, drawIndirectFirstInstance and drawIndirectCount where present.
- Buffers carry storage and indirect usage; a GPU-owned buffer is never swapped under a draw.
- Compute programs from shaders/compute.list (binding 0 parameters, 1.. storage buffers),
built by `ludic-dev shaders`; gpu_compute / gpu_dispatch / gpu_draw_mesh_indirect in gpu.ludic.
- R3D_VK_PROBE=1: a dispatch read back (OK on the RTX 3070 Ti).
- scatter_cull.comp: a tree layer's frustum test and LOD split on the GPU, with the lit, prepass,
impostor and shadow-LOD draws reading its records. Behind R3D_GPU_CULL=1 and off by default:
at the camp it is slower (43.0 fps against 53.3), because the frame's cost is per-draw
descriptor sets and it adds empty-level draws. Validation-clean; OpenGL frames unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
R3D_VK_PROF at the camp view (about 2950 draws a frame) showed the frame spent in bookkeeping:
19.5 ms a frame finding pipelines by string key and 16 ms building descriptor sets.
- Pipelines: a mesh carries an interned vertex-layout id (dropped only when an attribute's shape
changes, not when an instance buffer is swapped), render state packs into an int and the pass
formats into another; the program's last hit is tried first. The string path only builds.
- Samplers: each texture keeps the sampler for its parameters until they change.
- Descriptor sets: a program's last set is reused within the frame while its blocks and resolved
textures are unchanged; a uniform write that repeats the value it already holds changes nothing.
Headless on the RTX 3070 Ti at 1920x1080: 21.7 -> 53.3 fps (pipelines 0.2 ms, sets 8.5 ms, inside
draws 10 ms a frame; OpenGL 114 fps). The camp frame is unchanged and validation-clean. OpenGL frames
byte-identical at the five viewpoints; 59 self-tests pass; VKRES OK.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- A buffer re-uploaded after a draw this frame read it gets fresh storage, as OpenGL orphans
one; storage moved off or freed while the frame still reads it is destroyed after the frame's
submit. Draws mark the buffers they bind. No flush for buffer work.
- Mipmaps asked for mid-frame (the exposure measure, every frame) are recorded into the open frame
after its pass; growing a chain the first time keeps its one-shot path.
- R3D_VK_PROF prints, every 120 frames, draws and flushes per frame and the milliseconds spent
finding pipelines, filling descriptor sets and inside draws.
The gain was small - the camp view headless at 1920x1080 on the RTX 3070 Ti went from 21.3 to
21.7 fps (OpenGL: 114 fps) - so these flushes were not what holds the frame; the profile is how
the rest is found. The frame is unchanged and validation-clean; OpenGL frames byte-identical at the
five viewpoints with 59 self-tests passing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- gpu_select takes Vulkan for a window on Windows; gvk_init asks for VK_KHR_surface,
VK_KHR_win32_surface and VK_KHR_swapchain when the renderer will present. A window elsewhere
stays on OpenGL with the reason.
- gvk_open opens the window without a GL context and makes the screen images at its client area;
gvk_swap_make builds the swapchain on the Win32 surface (FIFO with vsync, mailbox or immediate
without) and rebuilds it when it is out of date, suboptimal or the window changes size.
- gvk_present blits the screen image into the acquired swapchain image, flipped (the screen keeps
OpenGL's bottom-up rows), and presents it.
- Samplers pointed at a texture unit with u_i and then fed by binding units - the actors do this -
read that unit's texture; on Vulkan they drew with the white stand-in (the tent, the log, the
chair, the workbench).
- gl.ll carries a weak win_gl_drawable for headless macOS builds.
On the RTX 3070 Ti the windowed build runs the valley through Vulkan at 3840x2160 with the HUD and
the frame-rate readout (16 fps: every upload still waits on the frame, and buffers are
host-visible). The headless Vulkan frame is unchanged; OpenGL frames byte-identical at the five
viewpoints with 59 self-tests passing. The actor fix is compiled and OpenGL-verified, not yet
seen on the PC; blades at the grass's middle distance still draw black on Vulkan.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
OpenGL links a vertex output to a fragment input by name; SPIR-V links them by location, and
glslang's --auto-map-locations numbered each stage in its own declaration order. The foliage
prepass's depth.frag declares v_wpos then v_uv where model.vert writes v_wpos, v_nrm, v_uv, so
the prepass read a normal as its texture coordinate, its alpha test cut every flower head and
leaf, and the lit pass (depth EQUAL) drew nothing over them. `ludic-dev shaders` now gives both
stages explicit locations: the vertex stage's out order numbers them and the fragment stage looks
each in up by name. All 45 variants checked: every fragment input sits on its vertex output.
Also:
- A clear still waiting for its pass when the framebuffer changes now runs on that framebuffer,
instead of becoming the load op of whichever pass began next.
- R3D_DUMP_ATLAS writes every impostor and card atlas a run bakes (build/atlas_<n>_*.ppm); the
40 baked on Vulkan match OpenGL's.
The PC's Vulkan frame now shows the flowers as OpenGL does, validation-clean. ludic-dev test 140
passed; OpenGL frames byte-identical at the five viewpoints; 59 self-tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three things the first Vulkan frames on the RTX 3070 Ti showed against OpenGL on the same PC:
- Exposure: the adaptation pass reads the HDR scene's smallest mip, and a render target made
without pixels had one level, so exposure came from a single texel. A target asked for mipmaps
now grows a full chain (level 0 kept), and passes draw through level-0 views keyed by the
image's generation.
- Foliage: the depth prepass and the lit pass (depth EQUAL) are different variants. Vulkan vertex
stages now declare an invariant gl_Position so both land on the same depth.
- Alpha to coverage is enabled only on a multisampled pass. OpenGL ignores it without MSAA; Vulkan
with one sample dropped every fragment under half alpha.
vk_resources also checks a big-endian 16-bit RGB upload (a normal map). VKRES OK; ludic-dev test
140 passed; the PC's Vulkan frame is validation-clean; OpenGL frames byte-identical at the five
viewpoints with 59 self-tests passing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
gpu.ludic's calls branch to the Vulkan backend when it was chosen and came up; every OpenGL
statement is unchanged, only guarded. R3D_GFX=vk selects it in a headless run (the window's
swapchain is the next milestone) and falls back to OpenGL, with the reason, when the device or
the SPIR-V manifest is missing.
- Render state is cached as before and turned into pipelines at the draw; u_* and sampler binds
go into the variant's uniform blocks; meshes, buffers and textures are gvk_* objects;
framebuffer binds are dynamic-rendering passes; same-size blits are image copies; the screen,
the photograph read-back, the present and the screenshot go through the frame.
- Work that submits on its own (uploads, read-backs, new or freed images and buffers) flushes the
frame first, so it runs in OpenGL's order. A read may take fewer channels than the image has
(the height field's R from its RGBA32F bake). Pipeline keys name vertex bindings by order, not
buffer handle, so re-pointed instance buffers keep their pipeline.
The valley renders at frame 90 validation-clean on the RTX 3070 Ti and on MoltenVK. OpenGL frames
byte-identical at the five viewpoints; 59 self-tests pass with no GL error.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The drawing half of the Vulkan backend, compiled into render3d and not yet reached from it:
- gvk_program: a program handle as its manifest variant - two SPIR-V modules, a descriptor set
layout (the vertex block at 0, the fragment block at 1, the samplers at their manifest bindings)
and a pipeline layout. Nothing is compiled at run time.
- gvk_pipeline: one pipeline per program, recorded vertex layout, render state and pass formats,
built the first time that combination draws. Attributes the shader does not read are left out;
the front face is clockwise, since neither API flips y between clip space and its target rows.
- gvk_uniform / gvk_u_set / gvk_bind_texture: loose uniforms written into each stage's block at
the manifest's offsets and array strides; gvk_draw_set copies the blocks into a per-frame ring
at the device's alignment and fills a descriptor set from a per-frame pool, with a white 1x1
texture for a sampler nothing was bound to.
- The frame: one command buffer; a framebuffer bind ends the pass and the next begins at its first
clear or draw (a clear that comes first is the load op); attachments move to attachment layouts
for the pass and back to SHADER_READ_ONLY after it, a cascade drawn through a view of its layer.
gvk_present submits and waits; gvk_screenshot reads the screen image back bottom row first.
r3d.ludic now imports gpu_manifest.ludic too. Textures remember their size.
OpenGL frames byte-identical at the five viewpoints; 59 self-tests pass; VKRES OK and VKDEVICE OK.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>