render3d: actors cast only into the cascades they can shade; Vulkan GPU timings and a mean frame split in R3D_PROF
Where the frame goes, before optimising it further. R3D_PROF now prints the average frame: CPU before the swap (the game's share apart), GPU and swap, and the renderer's CPU phases in frame order - the shadow pass split into cascade fit, scatter casters and actor casters. On Vulkan the per-pass GPU table comes from timestamp queries (host query reset asked for where the device has it); MoltenVK's attribution is tile-based and not to be trusted per pass, the PC's is. What it showed on the PC (camp): 4.3 ms CPU and 5.3 ms GPU a frame; the shadow pass was the largest CPU phase (1.8 ms) and actors half of that. Every actor within 300 m was drawn into all five cascades, though the outer two only shade receivers from 212 and 935 m out: 160 actors and 300 draws into each. cast_band_reaches - the flowers' reach test, now shared - skips an actor for a cascade it cannot shade (receivers counted from 0.85 of the previous split, where sunShadow's cross-fade begins). PC camp, two runs each: mean frame 9553/9601 -> 8664/8656 us; CPU 4.3 -> 3.7 ms; shadow GPU 1.14 -> 0.79 ms; actor-shadow CPU 1.02 -> 0.60 ms; 2039 -> 1411 draws; self-tests 59/59, validation 0. Mac: OpenGL shot viewpoints and the camp byte-identical; town 19 px at <= 2/255 on two flower stems a few metres from the camera - the accepted leftover-binding difference, no shadow; self-tests 59/59 on OpenGL and Vulkan. R3D_CAST_ALL=1 draws every caster into every cascade. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
ef745a141d
commit
5b9403c343
6 changed files with 134 additions and 10 deletions
|
|
@ -216,12 +216,12 @@ function gpu_program_free(p: int) -> void {
|
|||
}
|
||||
|
||||
# ---- GPU timers (R3D_PROF) ----------------------------------------------------------
|
||||
function gpu_query_new(n: int, ids: words) -> void { if gpu_kind == GPU_VK { return }; gl_gen_queries(n, ids) }
|
||||
function gpu_query_begin(id: int) -> void { if gpu_kind == GPU_VK { return }; gl_begin_query(GL_TIME_ELAPSED, id) }
|
||||
function gpu_query_end() -> void { if gpu_kind == GPU_VK { return }; gl_end_query(GL_TIME_ELAPSED) }
|
||||
function gpu_query_new(n: int, ids: words) -> void { if gpu_kind == GPU_VK { gvk_query_new(n, ids); return }; gl_gen_queries(n, ids) }
|
||||
function gpu_query_begin(id: int) -> void { if gpu_kind == GPU_VK { gvk_query_begin(id); return }; gl_begin_query(GL_TIME_ELAPSED, id) }
|
||||
function gpu_query_end() -> void { if gpu_kind == GPU_VK { gvk_query_end(); return }; gl_end_query(GL_TIME_ELAPSED) }
|
||||
# true once the query has its result; the nanoseconds (low 32 bits) are then in out[0]
|
||||
function gpu_query_result(id: int, out: words) -> bool {
|
||||
if gpu_kind == GPU_VK { return false }
|
||||
if gpu_kind == GPU_VK { return gvk_query_result(id, out) }
|
||||
gl_get_query_objectiv(id, GL_QUERY_RESULT_AVAILABLE, out)
|
||||
if out[0] == 0 { return false }
|
||||
gl_get_query_objectui64v(id, GL_QUERY_RESULT, out)
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue