feat(gl): OpenGL 4.1 and the ludic.render3d renderer
Some checks failed
ci / build-and-test (push) Waiting to run
commit-lint / conventional-commits (push) Waiting to run
bootstrap / cfree-fixpoint (push) Has been cancelled
docs / build-and-deploy (push) Successful in 34s

`Gl.*` binds the whole OpenGL 4.1 core API — every entry point of the
platform gl3.h with every GL_* constant, generated by `ludic-dev glgen`
with per-call ABI thunks. Windowed builds get an NSOpenGLContext on the
existing window at Retina resolution; headless builds render into an
offscreen CGL context, so a program that uses Gl.* renders and
screenshots identically under the test harness. It links gl.ll, the
thunks and OpenGL.framework only when used; every other build stays
byte-identical.

packages/ludic.render3d is a physically based renderer written on that
surface: HDRI image-based lighting, GPU-generated terrain with scanned
PBR materials, CDLOD, cascaded shadows, glTF with skinning, instanced
vegetation with impostors, procedural grass, water, SSAO, and an HDR
pipeline with bloom, auto-exposure and ACES.

It also carries this session's work on it: the terrain at half its cost
(10.3 -> 5.4 ms of frame), the streaming hitch that got worse the longer
you played, a resize that emptied the world, and the packaging that lets
a game use the renderer from its own repository — `ludic assets`, the
material manifest shipping with the package, and shader lookup falling
back to the install root. See changes/ for each, with its numbers.

The camping game that drove all of it has moved out to its own
repository, Maroon Lake; examples/rendering/smooth.ludic stays as the
renderer's example here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
Orkun ÇAKILKAYA 2026-09-10 03:31:12 +03:00
parent 470971bf70
commit f25289db20
90 changed files with 35316 additions and 19853 deletions

7
.gitignore vendored
View file

@ -47,3 +47,10 @@ __pycache__/
# archive lands in the root and a blanket `git add -A` will commit it.
*.tar.gz
*.tgz
# The CC0 Poly Haven downloads are fetched, not committed (`ludic-dev fetch-assets`
# reads the manifest that ships with the renderer, packages/ludic.render3d/assets.manifest,
# so a game outside this repository fetches the same set with `ludic assets`).
assets/polyhaven/hdri/
assets/polyhaven/textures/
assets/polyhaven/models/

View file

@ -1013,6 +1013,9 @@ literal; test any pointer/record/slice with `x == null` / `x != null` (an unset
# control quit() print(x) (a value + newline)
# convert str(x) -> str (int/bool/fixed -> text)
# length len(x) -> int (elements of a slice, or bytes of a string)
# OpenGL Gl.<snake_name>(…) every OpenGL 4.1 core entry point (glBindBuffer -> Gl.bind_buffer,
# GL_* constants as-is) float/double parameters take fixed; buffers are bytes/words
# Gl.open(width,height,title) Gl.swap() Gl.screenshot(path) Gl.program(vs,fs) Gl.vao() Gl.floats(n) …
# process arg_count()->int arg(i)->str (the command line; argv[0] included)
# exit(code) run(cmd) getenv(name) read_char()->int
# file_stderr()->ptr file_stdout()->ptr (handles for file_write)

View file

@ -0,0 +1,3 @@
Everything under assets/polyhaven/ is fetched from https://polyhaven.com and is
released by Poly Haven under CC0 1.0 (public domain). Files are not committed;
see manifest.txt for the exact sources. Re-fetch with tools/glgen/fetch_assets.sh.

28
changes/gl-render3d.md Normal file
View file

@ -0,0 +1,28 @@
bump: minor
type: feat
**OpenGL for Ludic, and a 3D renderer on it.** `Gl.*` binds the whole OpenGL 4.1
core API — every `gl*` entry point of the platform `gl3.h` as `Gl.<snake_name>(…)`
with every `GL_*` constant, generated by `ludic-dev glgen` with per-call ABI thunks
(`runtime/native/gl_thunks.ll`; float/double parameters take `fixed`). Windowed
builds get an `NSOpenGLContext` on the existing window at Retina resolution
(`cocoa.ll`); headless builds render into an offscreen CGL context, so a program
that uses `Gl.*` renders and screenshots identically under the test harness.
`Gl.open / swap / screenshot / program / vao / floats …` cover the glue, and the
IEEE-float helpers (`f_add`, `mem_put_f32`, …) let Q16.16 programs fill real
float vertex and uniform data. `Gl.*` links `gl.ll + gl_thunks.ll + OpenGL.framework`
only when used; every other build is byte-identical.
The `ludic.render3d` package (`packages/ludic.render3d`) is a physically based
renderer written on `Gl.*`: HDRI sky with image-based lighting (irradiance, GGX
prefiltered, split-sum BRDF, sun extracted from the map), GPU-generated terrain
with scanned PBR materials (stochastic anti-tiling, triplanar rock, slope/altitude
splatting), cascaded shadow maps with PCF and world-unit biasing, a glTF loader
for scanned models, instanced vegetation with baked impostors, procedural grass
and lupines with wind and translucency, 4x MSAA with alpha-to-coverage, SSAO,
still water, cloud shadows, aerial perspective, an HDR pipeline with bloom,
auto-exposure, ACES tonemapping, grading, sharpening and grain. See
`examples/rendering/smooth.ludic`, and the Maroon Lake game (git.workshopsoft.io/workshopsoft/maroon-lake) for a
game built on it. The renderer's CC0 materials are fetched with `ludic assets`.
**16-bit PNGs**: the renderer's texture loader keeps 16-bit samples (normal /
displacement maps) and uploads them as `RGB16` / `R16`.

View file

@ -0,0 +1,39 @@
bump: minor
type: feat
**A game can use the renderer from outside this repository.** `ludic.render3d` reads two
things from disk at run time — its GLSL, and the scanned CC0 materials — and both were
found only by a path relative to the working directory, so the renderer worked in a
Ludic checkout and nowhere else. A game living in its own repository now needs to copy
neither.
**The shaders come from the package**, wherever the package is. They belong to
`ludic.render3d` and ship with it, so the renderer looks for them beside the project
first (a Ludic checkout, where they are under `packages/`) and then under the install
root, `$LUDIC_HOME/packages/ludic.render3d` — the same place the compiler already
resolves `import "ludic.render3d/r3d.ludic"` from. Nothing to vendor, and no version of
the shaders that can drift from the version of the code that compiles them.
**`ludic assets [--force]`** fetches the scanned materials and the HDRI sky into
`assets/polyhaven/` of whatever project you run it in. The list of what to fetch is the
renderer's own — the renderer decides which materials it wants — so it moved out of the
repository's `assets/` and into the package as
`packages/ludic.render3d/assets.manifest`, where it ships with the toolchain. A game
does not keep its own copy of that list and so cannot fall out of step with the
renderer's material set. `ludic dev fetch-assets` is the same command from a checkout.
**A URL-shaped module built its binary into directories.** `project_name` took
everything after the last dot of the manifest's module path, which for a package
identified the way the package manager identifies them —
`git.host.io/user/name` — is inside the *host*: the build wrote
`build/io/user/name` instead of `build/name`. The last path segment comes first now, and
a dot inside that segment still separates namespace from package, so `ludic.render3d`
still builds as `render3d`.
**The camping game has moved out** to its own repository —
[Maroon Lake](https://git.workshopsoft.io/workshopsoft/maroon-lake) — taking
`examples/rendering/valley.ludic`, `hiker.ludic`, `camp/`, and 175 MB of survey data and
scanned kit with it. It was here as a demo of the renderer and became a game, and an
engine repository should not be carrying a game's assets. It is now the first consumer
of everything above, which is the point: what the renderer needs a game to be able to do,
it can now do from outside. `examples/rendering/smooth.ludic` stays as the renderer's
example in this tree.

View file

@ -0,0 +1,27 @@
bump: patch
type: fix
**Resizing the window (or entering fullscreen) no longer empties the world.** It left
`gl error 1286` — `GL_INVALID_FRAMEBUFFER_OPERATION` — on every frame from there on, with
the terrain, the trees, the grass and the water gone and only the sky drawn.
The sun-visibility pass added in `changes/terrain-perf.md` borrows the depth buffer the
frame is about to be drawn with, so that rasterising it doubles as a depth prepass. It
was storing that borrowed texture in its `Target`, and a `Target` deletes whatever its
`depth` names when it is freed. On the first resize the sequence was: `post_free` deletes
the frame's depth texture, `post_init` immediately makes the replacement — and GL hands
back the name that was just freed — and then the visibility target, rebuilt for the new
size, deleted that name believing it was its own. The scene framebuffer lost its depth
attachment. What is left is a colour-only framebuffer, which is *complete*, so drawing
carried on with no depth test at all: the sky is a fullscreen quad drawn last, and with
nothing left to fail against it painted over the entire valley. The 1286s came from the
passes whose own attachment now named a texture that no longer existed.
A borrowed attachment is never written into the target now, and the frame's depth is
attached afresh at the start of each pass — it is a different texture every time the
screen-sized buffers are rebuilt, and one `glFramebufferTexture2D` per pass is cheaper
than any scheme for noticing that it changed.
`R3D_RESIZE_AT=<frame>` rebuilds every screen-sized buffer from that frame on, cycling
through four drawable sizes every few frames. A window cannot be resized in a headless
run, so this is the only way to reach the path; it reproduced the fault in one frame and
now runs twenty resizes, with the fly camera and with the game, without an error.

36
changes/stream-hitch.md Normal file
View file

@ -0,0 +1,36 @@
bump: patch
type: perf
**Ground cover stops re-growing itself.** Walking a streamed world hitched, and the
hitch got worse the longer you played. Measured in the Maroon Lake game, with a new hitch
report rather than guessed at.
**The chunk cache had a cliff, not a slope.** A stream cached 4096 chunks and then
stopped remembering: past that the chunk was generated, used for one frame and thrown
away, so every ring walk regenerated it, for the rest of the session. It arrives after
enough of the map has been walked — six evictions' worth over seven kilometres, so an
ordinary session reaches it — and it is the point where cover starts visibly re-growing
as you turn. `stream_evict` now drops the half of the cache nobody has asked for in the
longest time (chunks carry the walk that last wanted them) and rebuilds the index over
what is left. Over a 7 km traversal: generation total **9073 ms → 2230 ms**, the worst
single frame's generation **11.4 ms → 3.1 ms**, median frame 10.8 → 9.0 ms. With a cache
deliberately sized to saturate early, the same run goes from 2748 frames generating to
1548, and from a 13.3 ms median to 9.2. `R3D_NOEVICT` restores the old behaviour for
comparison, `R3D_STREAM_CAP=<n>` sets the cache size.
The other half of that hitch was in the game's own cover generator, and went with it to
the Maroon Lake game's repository: its candidates were paying for a second noise field, four
height samples and a path distance before the drift field that rules out most of the
meadow — 2341 µs → 518 µs per chunk, bit-identical output. Worth repeating in any
generator: a `stream_fill` is called for tens of thousands of candidates per chunk, so
the order of its tests is most of its cost.
**The hitch report** (`R3D_PROF=1`) is what found both. It prints the slowest frames of
the run with what was in each: CPU versus GPU wait, cover generated, instance bytes
uploaded, the game's own tick, and the renderer phase that took longest. Alongside it,
per-chunk generation cost by stream and band, a census of what the caches hold, and a
stutter figure — the frame time a run spent beyond 1.2x its own median — because a mean
cannot show a hitch and a maximum is one unlucky frame.
It also found two content bugs in the game it was measured on, which is the report doing
its job: a cover stream whose placement rule never fires still pays full generation cost,
and the census makes that visible — 3364 cached chunks holding zero instances.

54
changes/terrain-perf.md Normal file
View file

@ -0,0 +1,54 @@
bump: patch
type: perf
**The ground costs half what it did.** Measured in the Maroon Lake game, the terrain was 10.3 ms of
a 22.2 ms frame; it is now 5.4 ms of 16.4 ms — 45 fps to 61 fps at 1080p, with the
frame otherwise unchanged (every viewpoint tested stays above 54 dB PSNR against the
old renderer, with no channel differing by more than 7/255).
**Measure by frame time, not by the pass timers.** `R3D_PROF`'s per-pass
`GL_TIME_ELAPSED` queries cannot be trusted on this driver: with the ground's shading
work removed the terrain query fell from 10.5 ms to 1.3 ms while the frame time did
not move at all. Every number above and below is a median real frame time, taken by
switching one thing off (`prof_ft_report`); the pass timers are still printed, and are
still useful for spotting a pass that appears out of nowhere, but they cannot size one.
`R3D_NOTERRAIN` skips the ground, `R3D_TNEARONLY` / `R3D_TFARONLY` draw every patch
with one tier's program, and `R3D_RES=<w>x<h>` renders at another size — the three
switches that say whether a cost is the ground, which tier it is in, and whether it is
pixels at all.
**Each detail tier is its own program.** `terrain.frag` holds a detailed near tier and
a cheap far one and chose between them per pixel, so every pixel of the valley walls
was compiled — and scheduled — for a near path it never ran. CDLOD selection now knows
which tiers a patch can contain: one that never comes within the split draws with
`FAR_ONLY`, one wholly inside it with `NEAR_ONLY`, and only the ring of patches that
straddle the band needs the program that holds both and cross-fades. Pixel-identical,
and it makes the tiers separately measurable: the near tier costs 7.5 ms over a whole
frame, the far tier 2.3 ms.
**The sun visibility is its own pass** (`tersun.frag`). The same CDLOD patches are
rasterised once into a screen-sized R8 buffer that holds nothing but each ground
pixel's sun visibility, and `terrain.frag` fetches it by fragment coordinate. The pass
costs 0.27 ms, shares the frame's depth buffer so it doubles as a depth prepass, and
takes the cascade read out of the shader that covers the screen. It picks its tier —
filtered PCF near, a single tap far — over the same cross-faded band the ground uses,
so the boundary is not a contour you can find on the hillside.
**Nothing is sampled for a weight of zero.** The ground sampled all four of its
materials for every pixel and then blended three of them at zero: a meadow pixel took
nine taps of triplanar rock, a cliff pixel nine taps of stochastic grass, and every
pixel in the valley took the snow tile and the four noise fields behind the lake's
shore wash — a wash that is a hairline along one shore within 120 m of the camera. The
survey photograph's classification now runs first, because it is what decides which
materials are present; each material block sits behind its own weight; the ridge field
that ragged the snow line is skipped 160 m below it, where it cannot change anything;
and inside the stochastic blend a cell's rotation, offset and rotated gradients are
computed inside its own test, so a cell whose sharpened weight rounds away costs
nothing. All of it exact where the weight is zero, and it is most of the win.
Material sampling is what remains (2.8 ms of the 5.4): scanned 2K tiles taken at 16x
anisotropy on ground seen at a grazing angle. `R3D_ANISO=<n>` sets the filter (the
default is unchanged at 16; 4 is worth 1.0 ms and 1 is worth 1.7 ms).
The shadow pass is 1.2 ms of the frame and has nothing to give: re-using the far
cascades between frames — their windows are snapped to a 14 m and a 64 m grid — is
worth 0.15 ms standing still and nothing while walking, so it is not in the tree.

View file

@ -0,0 +1,37 @@
---
id: gl
title: Gl
order: 60
---
OpenGL for Ludic. <code>Gl.*</code> binds the <strong>whole OpenGL 4.1 core API</strong>: every <code>gl*</code> entry point of the platform <code>gl3.h</code> is a method named by its snake case (<code>glBindBuffer</code> → <code>Gl.bind_buffer</code>, <code>glTexImage2D</code> → <code>Gl.tex_image2d</code>, <code>glDrawElementsInstanced</code> → <code>Gl.draw_elements_instanced</code>), with every <code>GL_*</code> constant available as written. Parameters keep the header's names as labels (<code>glVertexAttribPointer</code>'s <code>pointer</code> is called <code>offset</code>, <code>program</code> is <code>prog</code>); <code>GLfloat</code>/<code>GLdouble</code> parameters take <code>fixed</code>, pointer parameters take the raw <code>bytes</code>/<code>words</code> buffers Ludic already has, and 64-bit sizes take <code>long</code> (an <code>int</code> widens). The binding is generated by <code>ludic-dev glgen</code> from the header, with one ABI thunk per entry point.
A context comes from <code>Gl.open(width, height, title)</code>: windowed, an <code>NSOpenGLContext</code> on the game's window at the display's backing resolution (2× on Retina — <code>Gl.width()</code>/<code>Gl.height()</code> are the drawable's pixels); headless, an offscreen context with a framebuffer standing in for the screen (<code>Gl.screen_fbo()</code>), so the same program renders and screenshots under the test harness. <code>Gl.swap()</code> presents, <code>Gl.screenshot(path)</code> writes the screen as a PPM, <code>Gl.check(tag)</code> prints any pending error.
The glue: <code>Gl.program(vs, fs)</code> / <code>Gl.program5(vs, tcs, tes, gs, fs)</code> compile and link (logs on failure), <code>Gl.uniform(prog, name)</code>, <code>Gl.vao()</code>, <code>Gl.buffer()</code>, <code>Gl.texture()</code>, <code>Gl.framebuffer()</code> allocate, and float data is filled from Q16.16 with <code>Gl.floats(n)</code> / <code>Gl.put(buf, i, v)</code> / <code>Gl.bytes_of(n)</code> or from IEEE bits with <code>Gl.put_bits</code>; <code>Gl.f32(v)</code> and <code>Gl.fixed(bits)</code> convert single values. The bare <code>f_add</code>/<code>f_mul</code>/… helpers do IEEE single-precision arithmetic on those bit patterns for programs that want real floats on the CPU.
Using <code>Gl.*</code> links <code>gl.ll</code>, the thunks and <code>OpenGL.framework</code>; a program that does not is byte-identical to before. The <code>ludic.render3d</code> package is a physically based 3D renderer written on this surface (see <code>examples/rendering/smooth.ludic</code> here, and <a href="https://git.workshopsoft.io/workshopsoft/maroon-lake">Maroon Lake</a> for a game built on it).
```ludic
program Triangle {
property Marker { on: int = 1 }
model Anchor { Marker }
var prog: int = 0
var vao: int = 0
handler Boot phase Start {
spawn Anchor {}
if not Gl.open(width: 640, height: 360, title: "GL") { quit() }
prog = Gl.program(vs: "#version 410 core\nvoid main(){ gl_Position = vec4(float(gl_VertexID == 1) * 2.0 - 0.5, float(gl_VertexID == 2) * 2.0 - 0.5, 0.0, 1.0); }\n",
fs: "#version 410 core\nout vec4 o; void main(){ o = vec4(1.0, 0.5, 0.2, 1.0); }\n")
vao = Gl.vao()
}
handler Draw phase Render {
Gl.clear_color(red: 0.1, green: 0.1, blue: 0.15, alpha: 1.0)
Gl.clear(mask: GL_COLOR_BUFFER_BIT)
Gl.use_program(prog: prog)
Gl.bind_vertex_array(array: vao)
Gl.draw_arrays(mode: GL_TRIANGLES, first: 0, count: 3)
Gl.swap()
}
}
```

View file

@ -0,0 +1,49 @@
# gl_triangle.ludic — the smallest Gl.* program: one shaded triangle through a
# real OpenGL 4.1 core context. Windowed it draws to the screen each frame;
# headless it renders into an offscreen framebuffer and screenshots it, so the
# test harness can look at the result without a display.
# bin/ludic build examples/rendering/gl_triangle.ludic --headless && echo q | build/gl_triangle_headless
program GlTriangle {
property Marker { on: int = 1 }
model Anchor { Marker }
var prog: int = 0
var vao: int = 0
var frame: int = 0
const VS: string = "#version 410 core\nlayout(location=0) in vec2 p; layout(location=1) in vec3 c; out vec3 vc;\nvoid main(){ vc = c; gl_Position = vec4(p, 0.0, 1.0); }\n"
const FS: string = "#version 410 core\nin vec3 vc; out vec4 o;\nvoid main(){ o = vec4(vc, 1.0); }\n"
handler Boot phase Start {
spawn Anchor {}
if not Gl.open(width: 640, height: 360, title: "Ludic GL") { print("no GL context"); quit() }
print(Gl.get_string(name: GL_VERSION))
prog = Gl.program(vs: VS, fs: FS)
vao = Gl.vao()
let vbo = Gl.buffer()
Gl.bind_buffer(target: GL_ARRAY_BUFFER, buffer: vbo)
let v = Gl.floats(15)
Gl.put(v, 0, -0.8); Gl.put(v, 1, -0.8); Gl.put(v, 2, 1.0); Gl.put(v, 3, 0.2); Gl.put(v, 4, 0.1)
Gl.put(v, 5, 0.8); Gl.put(v, 6, -0.8); Gl.put(v, 7, 0.1); Gl.put(v, 8, 1.0); Gl.put(v, 9, 0.2)
Gl.put(v, 10, 0.0); Gl.put(v, 11, 0.8); Gl.put(v, 12, 0.2); Gl.put(v, 13, 0.3); Gl.put(v, 14, 1.0)
Gl.buffer_data(target: GL_ARRAY_BUFFER, size: Gl.bytes_of(15), data: v, usage: GL_STATIC_DRAW)
Gl.enable_vertex_attrib_array(index: 0)
Gl.vertex_attrib_pointer(index: 0, size: 2, type: GL_FLOAT, normalized: 0, stride: 20, offset: null)
Gl.enable_vertex_attrib_array(index: 1)
Gl.vertex_attrib_pointer(index: 1, size: 3, type: GL_FLOAT, normalized: 0, stride: 20, offset: Gl.ptr(null, 8))
Gl.check(tag: "boot")
}
handler Draw phase Render {
Gl.bind_framebuffer(target: GL_FRAMEBUFFER, framebuffer: Gl.screen_fbo())
Gl.viewport(x: 0, y: 0, width: Gl.width(), height: Gl.height())
Gl.clear_color(red: 0.1, green: 0.12, blue: 0.18, alpha: 1.0)
Gl.clear(mask: GL_COLOR_BUFFER_BIT | GL_DEPTH_BUFFER_BIT)
Gl.use_program(prog: prog)
Gl.bind_vertex_array(array: vao)
Gl.draw_arrays(mode: GL_TRIANGLES, first: 0, count: 3)
frame += 1
if frame == 1 { Gl.screenshot(path: "build/gl_triangle.ppm") }
Gl.swap()
}
}

View file

@ -0,0 +1,189 @@
# smooth.ludic — a synthetic 2 km test ground for the terrain renderer.
#
# The valley scene stands on a real survey (Copernicus GLO-30 resampled to a 4 m grid),
# which carries its own resampling lattice and quantisation. This scene carries none:
# the height field is an analytic function (see SMOOTH in heightgen.frag), so anything
# that still looks like a grid here belongs to the renderer, not the data.
#
# bin/ludic build examples/rendering/smooth.ludic && ./build/smooth
program Smooth {
import "ludic.render3d/r3d.ludic"
property Marker { on: int = 1 }
model Anchor { Marker }
var frame: int = 0
var shot_at: int = 40
var walk: bool = false
var spin: bool = false
var l_blades: Layer = null
var l_grass_a: Layer = null
var l_grass_b: Layer = null
var s_blades: Stream = null
var s_cards_a: Stream = null
var s_cards_b: Stream = null
var l_trees: Layer = null
function rnd() -> int { return fr(rng_range(0, 9999), 10000) }
function rnd_range(a: int, b: int) -> int { return f_lerp(a, b, rnd()) }
function smooth(a: int, b: int, x: int) -> int {
let t = f_clamp(f_div(f_sub(x, a), f_sub(b, a)), F_ZERO, F_ONE)
return f_mul(f_mul(t, t), f_sub(fi(3), f_mul(F_TWO, t)))
}
function slope_at(x: int, z: int) -> int {
let e = F_TWO
let dx = f_sub(terrain_height(f_add(x, e), z), terrain_height(f_sub(x, e), z))
let dz = f_sub(terrain_height(x, f_add(z, e)), terrain_height(x, f_sub(z, e)))
let ny = f_div(fi(4), f_sqrt(f_add(f_add(f_mul(dx, dx), f_mul(dz, dz)), fi(16))))
return f_sub(F_ONE, ny)
}
function noise01(x: int, z: int, scale: fixed) -> int {
let n = Noise.fbm2(f_fx(x) * scale, f_fx(z) * scale, 4)
return fl(n * 0.5 + 0.5)
}
# grass only where the ground is grassy: gentle, above the water, below the mountain
function stream_fill(s: Stream, cx: int, cz: int, band: int) -> void {
seed((cx * 73856093) ^ (cz * 19349663) ^ (band * 83492791) ^ (s.kind * 2654435761))
let size = s.size
let x0 = f_mul(fi(cx), size); let z0 = f_mul(fi(cz), size)
# Candidate spacing. These are metres between attempts, so halving one quadruples the
# work and the instance count: the first pass here was dense enough to bury the scene
# and cost most of the frame. Blades are only placed close to the eye, where they read
# as individual grass; past that the cards carry the cover.
var step = fl(1.0)
if s.kind == 0 { if band == 0 { step = fl(0.22) } else if band == 1 { step = fl(0.45) } else { step = fl(1.1) } }
else { if band == 0 { step = fl(0.9) } else if band == 1 { step = fl(1.8) } else { step = fl(4.0) } }
var z = z0
while f_ls(z, f_add(z0, size)) {
var x = x0
while f_ls(x, f_add(x0, size)) {
let px = f_add(x, f_mul(rnd(), step))
let pz = f_add(z, f_mul(rnd(), step))
let h = terrain_height(px, pz)
# Grassy ground only: above the water, off the steep parts, below the rim. Every
# test is a smooth ramp — a hard height cut carves the cover into contour rings,
# because the cut lands on a line of constant elevation.
var keep = smooth(fl(0.2), fl(1.6), h) # out of the water
keep = f_mul(keep, smooth(fi(90), fi(45), h)) # below the mountain
keep = f_mul(keep, smooth(fl(0.42), fl(0.16), slope_at(px, pz)))
keep = f_mul(keep, f_add(fl(0.25), f_mul(fl(0.9), noise01(px, pz, 0.02))))
keep = f_mul(keep, fl(0.75))
var sc = rnd_range(fl(0.16), fl(0.4))
if s.kind != 0 { sc = f_mul(rnd_range(fl(1.5), fl(2.5)), f_add(F_ONE, f_mul(fl(0.5), fi(band)))) }
if f_ls(rnd(), keep) {
stream_emit(s, px, f_sub(h, fl(0.03)), pz, sc, f_mul(rnd(), f_mul(F_TWO, F_PI)), rnd(), rnd_range(fl(0.6), F_ONE))
}
x = f_add(x, step)
}
z = f_add(z, step)
}
}
function scene_draw() -> void { scatter_draw() }
function scene_draw_casters() -> void { scatter_draw_casters(shadow_cascade_vp(sh_cascade)) }
handler Boot phase Start {
spawn Anchor {}
# 2 km square, generated analytically
TERRAIN_HALF = 1000
ter_smooth = true
if not r3d_init(1920, 1080, "Smooth") { quit() }
walk = Os.has_env("R3D_WALK")
spin = Os.has_env("R3D_SPIN")
if Os.has_env("R3D_SHOT") { shot_at = Text.to_int(Os.env("R3D_SHOT")) }
var cx = fi(0); var cz = fi(260)
if Os.has_env("R3D_CAM_X") { cx = fi(Text.to_int(Os.env("R3D_CAM_X"))) }
if Os.has_env("R3D_CAM_Z") { cz = fi(Text.to_int(Os.env("R3D_CAM_Z"))) }
var ch = fl(1.8); var cp = f_neg(fi(3)); var cy = fi(180)
if Os.has_env("R3D_CAM_H") { ch = fi(Text.to_int(Os.env("R3D_CAM_H"))) }
if Os.has_env("R3D_CAM_PITCH") { cp = fi(Text.to_int(Os.env("R3D_CAM_PITCH"))) }
if Os.has_env("R3D_CAM_YAW") { cy = fi(Text.to_int(Os.env("R3D_CAM_YAW"))) }
cam_set(cx, f_add(terrain_height(cx, cz), ch), cz, cy, cp)
# a small pond in the middle of the meadow — the plane self-clips to the basin, so
# its extent only has to cover the hollow, not the map
water_init(fl(4.0), F_ZERO, F_ZERO, fi(200), fi(200))
sky_set_yaw(f_rad(fi(120)))
# grass: blades underfoot, cards beyond
let dir = r3d_assets + "/models/grass_medium_01"
let ga = gltf_load(dir, "grass_medium_01_1k.gltf", "grass_medium_01_tall_a_LOD0")
let gb = gltf_load(dir, "grass_medium_01_1k.gltf", "grass_medium_01_mid_a_LOD0")
if ga == null or gb == null { print("smooth: no grass models"); return }
# blades are the expensive layer: keep them near, and cap them low
l_blades = layer_new(model_blade(), 400000, true, fl(3.0), F_ZERO, fi(70))
l_blades.blade = true
l_grass_a = layer_cards(ga, 200000, fl(0.3), fi(260))
l_grass_b = layer_cards(gb, 200000, fl(0.3), fi(260))
v3_set(l_grass_a.tint, fl(0.95), F_ONE, fl(0.85))
v3_set(l_grass_b.tint, fl(0.9), fl(0.98), fl(0.8))
l_grass_a.rough = fl(1.6); l_grass_b.rough = fl(1.6)
s_blades = stream_new(l_blades, fi(8), fi(60), fi(12), fi(28), fi(60), fi(60))
s_cards_a = stream_new(l_grass_a, fi(32), fi(240), fi(45), fi(110), fi(240), fi(240))
s_cards_b = stream_new(l_grass_b, fi(32), fi(240), fi(45), fi(110), fi(240), fi(240))
s_blades.kind = 0; s_cards_a.kind = 1; s_cards_b.kind = 2
place_trees()
}
# scattered firs on the gentle ground, thinning toward the pond and the rim
function place_trees() -> void {
let model = gltf_load(r3d_assets + "/models/fir_tree_01", "fir_tree_01_1k.gltf", "fir_tree_01_c_LOD0")
if model == null { print("smooth: no fir model"); return }
l_trees = layer_new(model, 20000, false, F_ZERO, fi(90), F_ZERO)
layer_set_impostor(l_trees, impostor_bake(model, 12, 512, 1024))
v3_set(l_trees.tint, fl(0.9), F_ONE, fl(0.85))
seed(4242)
var z = f_neg(fi(900))
while f_ls(z, fi(900)) {
var x = f_neg(fi(900))
while f_ls(x, fi(900)) {
let px = f_add(x, f_mul(rnd(), fi(14))); let pz = f_add(z, f_mul(rnd(), fi(14)))
let h = terrain_height(px, pz)
var keep = smooth(fi(7), fi(14), h) # above the pond shore
keep = f_mul(keep, smooth(fi(120), fi(60), h)) # below the rim
keep = f_mul(keep, smooth(fl(0.45), fl(0.2), slope_at(px, pz)))
keep = f_mul(keep, smooth(fl(0.45), fl(0.75), noise01(px, pz, 0.004)))
if f_ls(rnd(), f_mul(keep, fl(0.5))) {
let sc = rnd_range(fl(0.7), fl(1.35))
layer_add(l_trees, px, f_sub(h, fl(0.2)), pz, sc, f_mul(rnd(), f_mul(F_TWO, F_PI)), rnd(), F_ZERO)
}
x = f_add(x, fi(14))
}
z = f_add(z, fi(14))
}
}
handler Fly phase Input {
if walk { cam_move(fl(1.0), F_ZERO, F_ZERO, fl(0.003), F_ZERO); return }
if spin { cam_move(F_ZERO, F_ZERO, F_ZERO, fl(0.02), F_ZERO); return }
if not is_windowed() { return }
Input.poll()
var fwd = F_ZERO; var side = F_ZERO; var up = F_ZERO
let speed = fl(0.6)
if Input.key_down(key: 'w') { fwd = speed }
if Input.key_down(key: 's') { fwd = f_neg(speed) }
if Input.key_down(key: 'd') { side = speed }
if Input.key_down(key: 'a') { side = f_neg(speed) }
if Input.key_down(key: 'e') { up = speed }
if Input.key_down(key: 'q') { up = f_neg(speed) }
var dyaw = F_ZERO; var dpitch = F_ZERO
if Input.mouse_down(button: 0) {
dyaw = f_mul(fi(Input.mouse_dx()), fl(-0.004))
dpitch = f_mul(fi(Input.mouse_dy()), fl(-0.004))
}
if fwd != 0 or side != 0 or up != 0 or dyaw != 0 or dpitch != 0 { cam_move(fwd, side, up, dyaw, dpitch) }
if Input.key_pressed(key: 'p') {
let gy = terrain_height(cam_pos[0], cam_pos[2])
print(`R3D_CAM_X={string(f_to_int(cam_pos[0]))} R3D_CAM_Z={string(f_to_int(cam_pos[2]))} R3D_CAM_H={string(f_to_int(f_sub(cam_pos[1], gy)))} R3D_CAM_YAW={string(f_to_int(f_mul(cam_yaw, f_div(fi(180), F_PI))))} R3D_CAM_PITCH={string(f_to_int(f_mul(cam_pitch, f_div(fi(180), F_PI))))}`)
}
}
handler Draw phase Render {
r3d_frame(fl(Time.elapsed()))
frame += 1
if frame == shot_at { Gl.screenshot(path: "build/smooth.ppm") }
Gl.swap()
}
}

View file

@ -0,0 +1,118 @@
# ludic.render3d — what is implemented, where, and how it compares
Every item below names the file (and function or shader) that implements it, so
the claim can be checked in the tree. Paths are relative to `packages/ludic.render3d`
unless stated; `Gl.*` lives in `runtime/native/`.
## The OpenGL surface (`Gl.*`)
- **Every OpenGL 4.1 core entry point** — 478 of 478 `GLAPI` declarations in the
platform `gl3.h`, generated by `ludic-dev glgen` (`tools/ludic-cli/glgen.ludic`) into
`runtime/native/gl_api.ludic` (`grep -c '^extern function gl_'` = 478) with one ABI
thunk each in `runtime/native/gl_thunks.ll`; **901 `GL_*` constants**. Float and
double parameters take `fixed`; 64-bit sizes take `long`; `GLboolean` returns are
widened. macOS exposes 4.1 core as its ceiling, so this is the whole API the platform has.
- Context on the window at Retina resolution (`runtime/native/cocoa.ll`,
`win_gl_attach`), offscreen CGL context for headless runs (`runtime/native/gl.ll`),
screenshots, shader/program helpers, IEEE-float helpers over Q16.16 (`runtime/native/gl.ludic`).
## Rendering features (numbered so they can be counted)
Lighting and materials
1. Cook-Torrance GGX specular, Smith height-correlated visibility, Schlick Fresnel — `shaders/lighting.glsl` `shade()`
2. Diffuse image-based lighting from a convolved irradiance map — `sky.ludic` `sky_precompute`, `shaders/ibl_irradiance.frag`
3. GGX-prefiltered specular environment, 6 roughness levels (importance sampled, mip-filtered) — `shaders/ibl_prefilter.frag`
4. Split-sum BRDF lookup table — `shaders/ibl_brdf.frag`
5. HDRI sky with the sun extracted from the map and its irradiance integrated from the clipped texels — `texture.ludic` `hdr_decode`, `sky.ludic`
6. Sky rotation with all convolutions rebuilt (time-of-azimuth control) — `sky.ludic` `sky_set_yaw`
7. Normal mapping with derivative-reconstructed tangent frames (models) and a heightfield tangent basis (terrain) — `shaders/model.frag` `cotangentFrame`, `shaders/terrain.frag`
8. Distant-grass carpet: the clump cards baked straight down into a tiling tile (alpha = coverage) and draped on the ground beyond the blade rings with stochastic anti-tiling, so the far meadow carries the near cover's clumps and gaps — `scatter.ludic` `carpet_bake`, `shaders/terrain.frag` `sampleCarpet` (replaced the parallax pass, whose sampler slot it took)
9. Stochastic (triangle-grid) anti-tiling texture sampling with rotated taps and `textureGrad` — `shaders/terrain.frag` `sampleMat`
10. Triplanar mapping for steep rock — `shaders/terrain.frag` `sampleTri`
11. Foliage translucency and wrapped diffuse (thin-leaf subsurface approximation) — `shaders/model.frag` `FOLIAGE`
12. Two-sided lighting for thin cards — `shaders/model.frag` `CARD`
13. Ground-bounce fill light (one bounce from the sunlit meadow onto undersides) — `shaders/lighting.glsl` `bounce`
14. Screen-space global illumination: one indirect diffuse bounce gathered from the previous anti-aliased frame, plus ambient occlusion, half resolution, depth-aware blur — `shaders/ssgi.frag`, `shaders/ssao_blur.frag`, `post.ludic` `post_ssao_pass`
15. Specular occlusion from ambient occlusion — `shaders/lighting.glsl`
16. Cloud shadows projected along the sun from an animated cloud layer — `shaders/lighting.glsl` `cloudShadow`
17. Regional vegetation colouring (aspen groves, dry ridges) and slope-aspect colouring — `shaders/lighting.glsl` `regionTint`, `shaders/terrain.frag`
Shadows
18. Cascaded shadow maps, 5 cascades (16 / 60 / 250 / 1100 / 6000 m) fitted to frustum-slice bounding spheres, texel-snapped, PCF radius widened on the far cascades — `shadow.ludic` `shadow_fit`
19. PCF with rotated Poisson taps and interleaved-gradient noise, normal-offset and world-unit slope bias, cascade edge blending — `shaders/lighting.glsl` `sunShadow`
20. Per-layer caster culling (grass in near cascades, coarse terrain meshes for far cascades) — `scatter.ludic` `scatter_draw_casters`, `terrain.ludic`
21. Sun-facing impostor cards in the shadow pass with mip-biased coverage — `shaders/impostor.vert`, `shaders/impostor.frag`
Geometry, terrain and world
22. GPU-generated terrain height field (ridged multifractal, fbm) — `shaders/heightgen.frag`
23. Real-world terrain from the Copernicus GLO-30 elevation model with slope-scaled fractal detail; a lake bed carved under the surface the model records — the survey window in the Maroon Lake game (git.workshopsoft.io/workshopsoft/maroon-lake), `shaders/heightgen.frag` `DEM`, `terrain.ludic` `terrain_use_dem`, `terrain_lake`
24. Satellite orthophoto drape (Sentinel-2 true colour) sampled bicubically, blended in past the middle ground, keeping the scanned materials' luminance grain plus three-scale procedural grain — the survey window in the Maroon Lake game (git.workshopsoft.io/workshopsoft/maroon-lake), `shaders/terrain.frag` `u_ortho`
25. Satellite-classified materials and placement: bare ground (cobble shore near, talus far), snow, meadow, dense conifer stands (tree density and forest-floor duff follow the photograph) — `shaders/terrain.frag`, `terrain.ludic` `terrain_ortho_green/scree/forest`
26. Terrain material splatting by slope, altitude and a track mask; height-map CPU read-back for placement queries — `shaders/terrain.frag`, `terrain.ludic` `terrain_height`
27. Snow by altitude and slope, gully snow, forest-floor material under stands — `shaders/terrain.frag`
28. glTF 2.0 loading of photogrammetry models by byte-range reads from the `.bin` — `gltf.ludic`
29. GPU instancing with divisor attributes and per-instance transforms — `scatter.ludic`, `shaders/model.vert`
30. Baked impostors (16-view atlas with coverage and normals, hull-normal blending, LOD switch) — `scatter.ludic` `impostor_bake`, `shaders/impostor.*`
31. Baked branch-card trees: wedge slices of a scanned crown (sectors × height bands) on radial cards over the real trunk mesh — `scatter.ludic` `impostor_bake_wedges`, `model_branch_cards`
32. Card atlases baked from scanned grass clumps and a dense procedural lupine, instanced as crossed cards — `scatter.ludic` `layer_cards`, `model_lupine_dense`
33. Procedural grass blades with wind, dryness and base occlusion; distant blades widen, lie flat to the ground normal and sink into the carpet at the ring's edge — `scatter.ludic` `model_blade`, `shaders/model.vert`, `shaders/model.frag` `BLADE`
34. Wind animation (height-weighted sway, per-instance phase) — `shaders/model.vert` `WIND`
35. Chunk-streamed ground cover with distance bands (blades in four rings, densest underfoot), deterministic per-chunk generation, nearest-first amortised over frames with an instance budget — `stream.ludic`
36. Multi-scale noise clustering for natural placement, forest density by elevation/slope/aspect, avalanche chutes, shoreline scrub — the Maroon Lake game (git.workshopsoft.io/workshopsoft/maroon-lake)
37. Level-of-detail by distance: full scan mesh → branch cards → single impostor — `scatter.ludic` `layer_update`
38. Texture edge dilation for cut-out atlases; 16-bit PNG decoding to RGB16/R16; RGBE HDR decoding — `texture.ludic`
39. Anisotropic filtering and mipmapping on all scanned maps — `texture.ludic` `tex_upload`
40. Still water: planar reflection pass through a reflection matrix with clip-below discard, depth-tinted body, soft shores, ripple normals, foam streaks, sun glitter; the surface writes depth so occlusion and temporal passes see it — `water.ludic`, `shaders/water.frag`
Anti-aliasing, camera and post-processing
41. Temporal anti-aliasing: Halton-jittered projection, depth reprojection, YCoCg neighbourhood clamping, history ping-pong — `camera.ludic` `cam_set_jitter`, `shaders/taa.frag`
42. Optional MSAA with alpha-to-coverage and per-mip alpha sharpening — `post.ludic`, `shaders/impostor.frag`
43. HDR pipeline (RGBA16F) with NaN/inf sanitising at every stage — `post.ludic`, `shaders/lighting.glsl` `sane`
44. Bloom: 13-tap downsample pyramid with soft threshold, tent upsample — `shaders/bloom_down.frag`, `shaders/bloom_up.frag`
45. Auto-exposure by mean luminance with temporal adaptation — `post.ludic` `post_measure`
46. ACES filmic tonemapping with log-space contrast — `shaders/tonemap.frag`
47. Colour grading: white balance, lift/gain, saturation — `shaders/tonemap.frag`
48. Vignette, dithering, film grain — `shaders/tonemap.frag`, `shaders/sharpen.frag`
49. Luma unsharp sharpening — `shaders/sharpen.frag`
50. Height-based exponential fog / aerial perspective from the prefiltered sky with sun inscatter — `shaders/lighting.glsl` `applyFog`
51. Specular anti-aliasing (highlight cap) and firefly prevention — `shaders/lighting.glsl`
52. Fly camera with mouse look — the Maroon Lake game (git.workshopsoft.io/workshopsoft/maroon-lake)
53. Profiling switches per stage and a frame benchmark — `render.ludic` `r3d_env_flags`
Characters and moving objects
54. glTF skins: node hierarchy, joints and inverse bind matrices loaded per model, four-influence GPU skinning, poses set per node in the model's frame — `skin.ludic`, `shaders/skin.vert`, `quat.ludic`
55. Actors: models placed and moved per frame (skinned or rigid), lit and cast into every cascade — `actor.ludic`
56. The drawn terrain height on the CPU (cubic B-spline of the texels) for things that stand on it — `terrain.ludic` `terrain_height_smooth`
57. A third-person character controller with a distance-driven procedural gait (walk / run / idle), slope and water limits, and a settling orbit camera — the Maroon Lake game (git.workshopsoft.io/workshopsoft/maroon-lake)
58. A 2D overlay over the finished frame: batched rectangles, images and proportional text from a baked font atlas (`tools/blender/font_build.py`) — `overlay.ludic`, `shaders/overlay.*`
59. Time of day over the HDRI: a solar arc drives the sun's direction and colour, the sky's light scales toward a deep-blue starry night, a moon lights the night, the visible sky turns with the sun and its convolutions rebake when far off, the height-field shadow rebakes as the sun moves, auto-exposure is capped at night — `daylight.ludic`, `shaders/sky.frag`, `shaders/lighting.glsl`
60. One point light (the campfire) in the shading — `shaders/lighting.glsl` `fireLight`
61. Static collision: ground circles in a cell grid with push-out and a segment test — `collide.ludic`
62. Cut-out and emissive actors (a flame's cards) — `actor.ludic`, `shaders/model.frag` `u_emissive`
63. Per-instance skin clones (a herd posed individually), per-part tints and hidden parts on actors, uniform-location caching with frustum and distance culling — `skin.ludic` `skin_clone`, `actor.ludic`
64. A second point / cone light (a torch or flashlight in the hand) — `daylight.ludic` `daylight_hand`, `shaders/lighting.glsl` `handLight`
65. Auto-exposure adapted on the GPU: a 1x1 pass eases last frame's value toward key / mean and the tonemapper samples it, so nothing is read back (a pixel-buffer read still synchronised on Apple's GL) — `post.ludic` `post_measure`, `shaders/adapt.frag`
66. A PPM reader with box-filtered downscale (photo thumbnails) and word-wrapped overlay text — `texture.ludic` `tex_load_ppm`, `overlay.ludic` `ov_text_wrap`
67. The overlay batches a whole frame into one upload and draws per texture range with optional scissor clips (per-change uploads stalled the driver: 40 ms -> 25 ms), plus nine-slice, rotated, line, disc and arc primitives — `overlay.ludic` `ov_flush`, `ov_clip`, `ov_nine`, `ov_sub_rot`, `ov_line`, `ov_disc`, `ov_arc`
68. Weather on the daylight: an overcast factor greys and dims the sun and the sky's light, a fog multiplier thickens the air, a lightning flash — `daylight.ludic` `daylight_weather`
## Against Unreal Engine 5 and RAGE (RDR2), honestly
| Technique | UE5 / RAGE | Here |
|---|---|---|
| Physically based shading, IBL, split-sum | both | implemented (1–4) |
| Cascaded shadow maps with PCF | both (UE5 also virtual shadow maps) | implemented (18–21); no virtual shadow maps |
| Temporal AA / upsampling | TSR / TAA | TAA implemented (41), no upscaling |
| Screen-space AO / GI | SSAO, SSGI, Lumen | SSGI + AO implemented (14); no Lumen-class GI, no ray tracing (not available on OpenGL 4.1) |
| Volumetric clouds and fog | both | HDRI sky with cloud shadows and analytic fog; no ray-marched clouds |
| Foliage: impostors, wind, translucency | both | implemented (30–34) |
| Nanite / virtualised geometry | UE5 | not applicable; LOD chain (37) instead |
| Terrain: height fields, layered materials, real-world data | both | implemented (22–27) |
| Water with planar reflection | RAGE planar, UE SSR | planar reflection implemented (40); no screen-space reflection |
| Post: bloom, exposure, tonemap, grading, DoF, motion blur | both | bloom/exposure/tonemap/grading implemented (44–49); no depth of field, no motion blur |
| Streaming world | both | chunk streaming of cover (35); terrain is one 8 km tile |
| Compute shaders, indirect draw, bindless | both | not available in OpenGL 4.1 core on macOS |
Where the remaining visual gap sits: photogrammetry cards for several meadow
species, ray-traced or probe-based global illumination, virtual shadow maps, and
volumetric clouds. Those are the techniques that separate this from the reference.

View file

@ -0,0 +1,204 @@
# ============================================================================
# actor.ludic — a model placed once and moved every frame (a character, a prop
# that animates): the counterpart of a scatter Layer for the handful of things
# that are not instanced. Skinned models pose through their Skin (skin.ludic);
# rigid ones take the same path with u_skinned = 0. The game draws them from
# scene_draw / scene_draw_casters through actor_draw / actor_draw_casters.
#
# Uniform locations are looked up once per program (glGetUniformLocation by
# name is slow on this driver, and a hundred actors over five cascades made it
# the frame's biggest CPU cost); the scene's lighting is bound once per program
# per pass, and actors outside the view or beyond their cull distance are skipped.
# ============================================================================
property Actor {
model: Model,
x: int = 0, # float bits, metres
y: int = 0,
z: int = 0,
yaw: int = 0, # radians; 0 faces -z like the camera
scale: int = 0,
tint: words,
rough: int = 0,
mat: words, # the model matrix
visible: bool = true,
casts: bool = true,
id: int = 0,
cutout: bool = false, # alpha-tested (a flame's cards)
emissive: int = 0, # float bits: self-lit strength
skin: Skin, # this instance's own pose (skin_clone); null: the model's
ptint: words, # 4 per primitive: on flag, r, g, b (a part's own colour)
hide: words, # 1 per primitive: skip it (a cap taken off)
radius: int = 0, # float bits: bounding radius for culling (0 = never culled)
cull: int = 0 # float bits: not drawn beyond this distance (0 = always)
}
# a program and its uniform locations
property AcProg {
prog: int = 0,
l_model: int = -1,
l_bones: int = -1,
l_skin: int = -1,
l_lvp: int = -1,
l_tint: int = -1,
l_rough: int = -1,
l_emis: int = -1,
l_view: int = -1,
l_proj: int = -1,
l_mh: int = -1,
scene_frame: int = -1 # the frame its scene uniforms were bound
}
var ac_lit: AcProg = null
var ac_lit_cut: AcProg = null
var ac_sh: AcProg = null
var ac_sh_cut: AcProg = null
var ac_actors: []Actor = null
var ac_next_id: int = 0
var ac_frame: int = 0
function ac_prog_new(vs: string, fs: string, defs: string) -> AcProg {
let a = new AcProg
a.prog = r3d_program(vs, fs, defs)
let p = a.prog
a.l_model = gl_uniform(p, "u_model"); a.l_skin = gl_uniform(p, "u_skinned"); a.l_lvp = gl_uniform(p, "u_light_vp")
a.l_bones = gl_uniform(p, "u_bones[0]"); if a.l_bones < 0 { a.l_bones = gl_uniform(p, "u_bones") }
a.l_tint = gl_uniform(p, "u_tint"); a.l_rough = gl_uniform(p, "u_rough_scale"); a.l_emis = gl_uniform(p, "u_emissive")
a.l_view = gl_uniform(p, "u_view"); a.l_proj = gl_uniform(p, "u_proj"); a.l_mh = gl_uniform(p, "u_model_h")
# the samplers never move: units 0, 1, 2
gl_use_program(p)
gl_uniform1i(gl_uniform(p, "u_diff"), 0); gl_uniform1i(gl_uniform(p, "u_nrm"), 1); gl_uniform1i(gl_uniform(p, "u_arm"), 2)
return a
}
function actor_init() -> void {
ac_lit = ac_prog_new("skin.vert", "model.frag", "")
ac_lit_cut = ac_prog_new("skin.vert", "model.frag", "#define ALPHA_TEST\n")
ac_sh = ac_prog_new("skin.vert", "shadow.frag", "#define SHADOW_PASS\n")
ac_sh_cut = ac_prog_new("skin.vert", "shadow.frag", "#define SHADOW_PASS\n#define ALPHA_TEST\n")
ac_actors = new []Actor
}
function actor_remove(a: Actor) -> void {
if ac_actors == null { return }
let keep = new []Actor
for i in 0 .. len(ac_actors) { if ac_actors[i].id != a.id { push(keep, ac_actors[i]) } }
ac_actors = keep
}
# colour one named part of the model (a material name from the file)
function actor_tint_part(a: Actor, name: string, r: int, g: int, b: int) -> void {
if a.model == null { return }
let n = len(a.model.prims)
if a.ptint == null { a.ptint = words(n * 4); for i in 0 .. n * 4 { a.ptint[i] = 0 } }
for i in 0 .. n { if a.model.prims[i].name == name { a.ptint[i * 4] = 1; a.ptint[i * 4 + 1] = r; a.ptint[i * 4 + 2] = g; a.ptint[i * 4 + 3] = b } }
}
function actor_hide_part(a: Actor, name: string, hidden: bool) -> void {
if a.model == null { return }
let n = len(a.model.prims)
if a.hide == null { a.hide = words(n); for i in 0 .. n { a.hide[i] = 0 } }
var v = 0
if hidden { v = 1 }
for i in 0 .. n { if a.model.prims[i].name == name { a.hide[i] = v } }
}
function actor_new(model: Model) -> Actor {
let a = new Actor
a.model = model
a.scale = F_ONE
a.tint = v3_new(F_ONE, F_ONE, F_ONE)
a.rough = F_ONE
a.mat = m4_new()
ac_next_id += 1; a.id = ac_next_id
if model != null { a.radius = f_add(f_max(model.radius, model.height), F_ONE) }
a.cull = fi(450)
if ac_actors == null { ac_actors = new []Actor }
push(ac_actors, a)
return a
}
function actor_place(a: Actor, x: int, y: int, z: int, yaw: int) -> void {
a.x = x; a.y = y; a.z = z; a.yaw = yaw
m4_trs(a.mat, x, y, z, yaw, a.scale)
}
# the scene's lighting for a lit program, once per frame
function ac_bind_scene(ap: AcProg) -> void {
if ap.scene_frame == ac_frame { return }
ap.scene_frame = ac_frame
let p = ap.prog
gl_use_program(p)
u_mat4(ap.l_view, cam_view)
u_mat4(ap.l_proj, cam_proj)
u_f(ap.l_mh, F_ZERO)
sky_bind_lighting(p)
shadow_bind(p)
fog_bind(p)
}
function ac_visible(a: Actor, shadow: bool) -> bool {
if not a.visible or a.model == null { return false }
if shadow and not a.casts { return false }
if a.cull != 0 {
let dx = f_sub(a.x, cam_pos[0]); let dz = f_sub(a.z, cam_pos[2])
let d2 = f_add(f_mul(dx, dx), f_mul(dz, dz))
var c = a.cull
if shadow { c = f_min(c, fi(300)) }
if f_gt(d2, f_mul(c, c)) { return false }
}
if not shadow and a.radius != 0 {
let r = f_mul(a.radius, a.scale)
if not cam_sphere_visible(a.x, f_add(a.y, r), a.z, f_mul(r, fl(1.5))) { return false }
}
return true
}
function actor_draw_one(a: Actor, ap: AcProg, shadow: bool) -> void {
let p = ap.prog
gl_use_program(p)
u_mat4(ap.l_model, a.mat)
var skinned = F_ZERO
if a.skin != null { skinned = F_ONE; gl_uniform_matrix4fv(ap.l_bones, a.skin.n_joints, 0, a.skin.bones) }
else if a.model.skin != null { skinned = F_ONE; gl_uniform_matrix4fv(ap.l_bones, a.model.skin.n_joints, 0, a.model.skin.bones) }
u_f(ap.l_skin, skinned)
if not shadow {
u_f(ap.l_emis, a.emissive)
u_f(ap.l_rough, a.rough)
}
gl_disable(GL_CULL_FACE)
let model = a.model
for i in 0 .. len(model.prims) {
let pr = model.prims[i]
if a.hide != null and a.hide[i] != 0 { continue }
if not shadow {
if a.ptint != null and a.ptint[i * 4] != 0 { u_f3(ap.l_tint, a.ptint[i * 4 + 1], a.ptint[i * 4 + 2], a.ptint[i * 4 + 3]) }
else { u_v3(ap.l_tint, a.tint) }
}
gl_active_texture(GL_TEXTURE0); gl_bind_texture(GL_TEXTURE_2D, pr.diff)
if not shadow {
gl_active_texture(GL_TEXTURE0 + 1); gl_bind_texture(GL_TEXTURE_2D, pr.nrm)
gl_active_texture(GL_TEXTURE0 + 2); gl_bind_texture(GL_TEXTURE_2D, pr.arm)
}
mesh_draw(pr.mesh)
}
}
function actor_draw() -> void {
if ac_actors == null { return }
ac_frame += 1
for i in 0 .. len(ac_actors) {
let a = ac_actors[i]
if not ac_visible(a, false) { continue }
var ap = ac_lit
if a.cutout { ap = ac_lit_cut }
ac_bind_scene(ap)
actor_draw_one(a, ap, false)
}
gl_enable(GL_CULL_FACE)
}
function actor_draw_casters(light_vp: words) -> void {
if ac_actors == null { return }
gl_use_program(ac_sh.prog); u_mat4(ac_sh.l_lvp, light_vp)
gl_use_program(ac_sh_cut.prog); u_mat4(ac_sh_cut.l_lvp, light_vp)
for i in 0 .. len(ac_actors) {
let a = ac_actors[i]
if not ac_visible(a, true) { continue }
var ap = ac_sh
if a.cutout { ap = ac_sh_cut }
actor_draw_one(a, ap, true)
}
}

View file

@ -0,0 +1,49 @@
hdri/kloofendal_48d_partly_cloudy_puresky_4k.hdr https://dl.polyhaven.org/file/ph-assets/HDRIs/hdr/4k/kloofendal_48d_partly_cloudy_puresky_4k.hdr
textures/aerial_grass_rock_diff_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/aerial_grass_rock/aerial_grass_rock_diff_2k.png
textures/aerial_grass_rock_nor_gl_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/aerial_grass_rock/aerial_grass_rock_nor_gl_2k.png
textures/aerial_grass_rock_arm_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/aerial_grass_rock/aerial_grass_rock_arm_2k.png
textures/aerial_grass_rock_disp_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/aerial_grass_rock/aerial_grass_rock_disp_2k.png
textures/grass_path_2_diff_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/grass_path_2/grass_path_2_diff_2k.png
textures/grass_path_2_nor_gl_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/grass_path_2/grass_path_2_nor_gl_2k.png
textures/grass_path_2_arm_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/grass_path_2/grass_path_2_arm_2k.png
textures/forest_leaves_04_diff_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/forest_leaves_04/forest_leaves_04_diff_2k.png
textures/forest_leaves_04_nor_gl_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/forest_leaves_04/forest_leaves_04_nor_gl_2k.png
textures/forest_leaves_04_arm_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/forest_leaves_04/forest_leaves_04_arm_2k.png
textures/gray_rocks_diff_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/gray_rocks/gray_rocks_diff_2k.png
textures/gray_rocks_nor_gl_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/gray_rocks/gray_rocks_nor_gl_2k.png
textures/gray_rocks_arm_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/gray_rocks/gray_rocks_arm_2k.png
textures/snow_02_diff_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/snow_02/snow_02_diff_2k.png
textures/snow_02_nor_gl_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/snow_02/snow_02_nor_gl_2k.png
textures/snow_02_arm_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/snow_02/snow_02_arm_2k.png
models/fir_tree_01/fir_tree_01_1k.gltf https://dl.polyhaven.org/file/ph-assets/Models/gltf/1k/fir_tree_01/fir_tree_01_1k.gltf
models/fir_tree_01/fir_tree_01.bin https://dl.polyhaven.org/file/ph-assets/Models/gltf/1k/fir_tree_01/fir_tree_01.bin
models/fir_tree_01/textures/fir_tree_01_bark_diff_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/fir_tree_01/fir_tree_01_bark_diff_1k.jpg
models/fir_tree_01/textures/fir_tree_01_bark_nor_gl_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/fir_tree_01/fir_tree_01_bark_nor_gl_1k.jpg
models/fir_tree_01/textures/fir_tree_01_bark_arm_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/fir_tree_01/fir_tree_01_bark_arm_1k.jpg
models/fir_tree_01/textures/fir_tree_01_trunk_a_diff_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/fir_tree_01/fir_tree_01_trunk_a_diff_1k.jpg
models/fir_tree_01/textures/fir_tree_01_trunk_a_nor_gl_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/fir_tree_01/fir_tree_01_trunk_a_nor_gl_1k.jpg
models/fir_tree_01/textures/fir_tree_01_trunk_a_arm_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/fir_tree_01/fir_tree_01_trunk_a_arm_1k.jpg
models/fir_tree_01/textures/fir_tree_01_trunk_b_diff_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/fir_tree_01/fir_tree_01_trunk_b_diff_1k.jpg
models/fir_tree_01/textures/fir_tree_01_trunk_b_nor_gl_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/fir_tree_01/fir_tree_01_trunk_b_nor_gl_1k.jpg
models/fir_tree_01/textures/fir_tree_01_trunk_b_arm_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/fir_tree_01/fir_tree_01_trunk_b_arm_1k.jpg
models/fir_tree_01/textures/fir_tree_01_twig_diff_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/fir_tree_01/fir_tree_01_twig_diff_1k.jpg
models/fir_tree_01/textures/fir_tree_01_twig_nor_gl_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/fir_tree_01/fir_tree_01_twig_nor_gl_1k.jpg
models/fir_tree_01/textures/fir_tree_01_twig_arm_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/fir_tree_01/fir_tree_01_twig_arm_1k.jpg
models/grass_medium_01/grass_medium_01_1k.gltf https://dl.polyhaven.org/file/ph-assets/Models/gltf/1k/grass_medium_01/grass_medium_01_1k.gltf
models/grass_medium_01/grass_medium_01.bin https://dl.polyhaven.org/file/ph-assets/Models/gltf/1k/grass_medium_01/grass_medium_01.bin
models/grass_medium_01/textures/grass_medium_01_diff_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/grass_medium_01/grass_medium_01_diff_1k.jpg
models/grass_medium_01/textures/grass_medium_01_nor_gl_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/grass_medium_01/grass_medium_01_nor_gl_1k.jpg
models/grass_medium_01/textures/grass_medium_01_arm_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/grass_medium_01/grass_medium_01_arm_1k.jpg
models/boulder_01/boulder_01_1k.gltf https://dl.polyhaven.org/file/ph-assets/Models/gltf/1k/boulder_01/boulder_01_1k.gltf
models/boulder_01/boulder_01.bin https://dl.polyhaven.org/file/ph-assets/Models/gltf/1k/boulder_01/boulder_01.bin
models/boulder_01/textures/boulder_01_diff_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/boulder_01/boulder_01_diff_1k.jpg
models/boulder_01/textures/boulder_01_nor_gl_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/boulder_01/boulder_01_nor_gl_1k.jpg
models/boulder_01/textures/boulder_01_arm_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/boulder_01/boulder_01_arm_1k.jpg
models/celandine_01/celandine_01_1k.gltf https://dl.polyhaven.org/file/ph-assets/Models/gltf/1k/celandine_01/celandine_01_1k.gltf
models/celandine_01/celandine_01.bin https://dl.polyhaven.org/file/ph-assets/Models/gltf/1k/celandine_01/celandine_01.bin
models/celandine_01/textures/celandine_01_diff_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/celandine_01/celandine_01_diff_1k.jpg
models/celandine_01/textures/celandine_01_nor_gl_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/celandine_01/celandine_01_nor_gl_1k.jpg
models/celandine_01/textures/celandine_01_arm_1k.jpg https://dl.polyhaven.org/file/ph-assets/Models/jpg/1k/celandine_01/celandine_01_arm_1k.jpg
textures/cliff_side_diff_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/cliff_side/cliff_side_diff_2k.png
textures/cliff_side_nor_gl_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/cliff_side/cliff_side_nor_gl_2k.png
textures/cliff_side_arm_2k.png https://dl.polyhaven.org/file/ph-assets/Textures/png/2k/cliff_side/cliff_side_arm_2k.png

View file

@ -0,0 +1,101 @@
# ============================================================================
# camera.ludic — a first-person fly camera (yaw / pitch, metres) and its view /
# projection matrices. Float bits throughout; see fmath.ludic.
# ============================================================================
var cam_pos: words = null # x, y, z
var cam_yaw: int = 0 # radians, 0 = looking down -z
var cam_pitch: int = 0
var cam_fov: int = 0 # vertical, radians
var cam_near: int = 0
var cam_far: int = 0
var cam_aspect: int = 0
var cam_view: words = null
var cam_proj: words = null
var cam_vp: words = null
var cam_inv_vp: words = null
var cam_inv_proj: words = null
var cam_vp_clean: words = null # view-projection, kept for depth reconstruction
var cam_inv_vp_clean: words = null
var cam_fwd: words = null
var cam_right: words = null
# the view frustum's four side planes (a, b, c, d), float bits: left, right, bottom, top;
# from the clean (unjittered) view-projection, column-major m[col * 4 + row]
var cam_planes: words = null
function cam_init(aspect: int) -> void {
cam_pos = v3_new(F_ZERO, fi(2), F_ZERO)
cam_view = m4_new(); cam_proj = m4_new(); cam_vp = m4_new(); cam_inv_vp = m4_new(); cam_inv_proj = m4_new()
cam_vp_clean = m4_new(); cam_inv_vp_clean = m4_new()
cam_fwd = v3_new(F_ZERO, F_ZERO, f_neg1())
cam_right = v3_new(F_ONE, F_ZERO, F_ZERO)
cam_fov = f_rad(fi(42))
cam_near = fl(0.3)
cam_far = fi(14000)
cam_aspect = aspect
cam_update()
}
function cam_begin_frame(n: int, w: int, h: int) -> void {
cam_update()
}
function cam_set(x: int, y: int, z: int, yaw_deg: int, pitch_deg: int) -> void {
v3_set(cam_pos, x, y, z)
cam_yaw = f_rad(yaw_deg)
cam_pitch = f_rad(pitch_deg)
cam_update()
}
function cam_update() -> void {
let cy = f_cos(cam_yaw); let sy = f_sin(cam_yaw)
let cp = f_cos(cam_pitch); let sp = f_sin(cam_pitch)
v3_set(cam_fwd, f_neg(f_mul(sy, cp)), sp, f_neg(f_mul(cy, cp)))
v3_set(cam_right, cy, F_ZERO, f_neg(sy))
let at = words(3)
v3_add(at, cam_pos, cam_fwd)
let up = v3_new(F_ZERO, F_ONE, F_ZERO)
m4_look_at(cam_view, cam_pos, at, up)
m4_perspective(cam_proj, cam_fov, cam_aspect, cam_near, cam_far)
m4_mul(cam_vp_clean, cam_proj, cam_view)
m4_inverse(cam_inv_vp_clean, cam_vp_clean)
m4_mul(cam_vp, cam_proj, cam_view)
m4_inverse(cam_inv_vp, cam_vp)
m4_inverse(cam_inv_proj, cam_proj)
free(at); free(up)
cam_planes_update()
}
function cam_planes_update() -> void {
if cam_planes == null { cam_planes = words(16) }
let m = cam_vp_clean
for p in 0 .. 4 {
var r = 0
if p >= 2 { r = 1 }
var sg = F_ONE
if p == 1 or p == 3 { sg = f_neg(F_ONE) }
var a = f_add(m[3], f_mul(sg, m[r]))
var b = f_add(m[7], f_mul(sg, m[4 + r]))
var c = f_add(m[11], f_mul(sg, m[8 + r]))
var d = f_add(m[15], f_mul(sg, m[12 + r]))
let inv = f_div(F_ONE, f_sqrt(f_add(f_add(f_mul(a, a), f_mul(b, b)), f_mul(c, c))))
cam_planes[p * 4] = f_mul(a, inv); cam_planes[p * 4 + 1] = f_mul(b, inv)
cam_planes[p * 4 + 2] = f_mul(c, inv); cam_planes[p * 4 + 3] = f_mul(d, inv)
}
}
# is a sphere (float bits) at least partly inside the side planes of the view?
function cam_sphere_visible(x: int, y: int, z: int, r: int) -> bool {
if cam_planes == null { return true }
let nr = f_neg(r)
for p in 0 .. 4 {
let o = p * 4
let dist = f_add(f_add(f_add(f_mul(cam_planes[o], x), f_mul(cam_planes[o + 1], y)), f_mul(cam_planes[o + 2], z)), cam_planes[o + 3])
if f_ls(dist, nr) { return false }
}
return true
}
# fly: forward/strafe in metres, turn in radians
function cam_move(fwd: int, strafe: int, up: int, dyaw: int, dpitch: int) -> void {
cam_yaw = f_add(cam_yaw, dyaw)
cam_pitch = f_clamp(f_add(cam_pitch, dpitch), f_neg(fl(1.5)), fl(1.5))
v3_madd(cam_pos, cam_pos, cam_fwd, fwd)
v3_madd(cam_pos, cam_pos, cam_right, strafe)
cam_pos[1] = f_add(cam_pos[1], up)
cam_update()
}

View file

@ -0,0 +1,99 @@
# ============================================================================
# collide.ludic — the static colliders of the world as circles on the ground
# plane (a trunk, a boulder, a tent), sorted once into 16 m cells over the whole
# terrain. A moving thing asks col_resolve for its position pushed out of every
# circle it overlaps: three by three cells, a few dozen tests, no broad phase
# needed. Float bits, metres.
# ============================================================================
const COL_CELL: int = 16
const COL_CAP: int = 120000
var col_x: words = null
var col_z: words = null
var col_r: words = null
var col_n: int = 0
var col_side: int = 0 # cells per side
var col_start: words = null # per cell: first index into col_sorted (side*side + 1)
var col_sorted: words = null
var col_built: bool = false
var col_out: words = null # the resolved position (x, z)
function col_add(x: int, z: int, r: int) -> void {
if col_x == null { col_x = words(COL_CAP); col_z = words(COL_CAP); col_r = words(COL_CAP); col_out = words(2) }
if col_n >= COL_CAP { return }
col_x[col_n] = x; col_z[col_n] = z; col_r[col_n] = r
col_n += 1
col_built = false
}
function col_cell_of(v: int, origin: int) -> int {
var c = f_to_int(f_floor(f_div(f_add(f_sub(v, origin), fi(TERRAIN_HALF)), fi(COL_CELL))))
if c < 0 { c = 0 }
if c > col_side - 1 { c = col_side - 1 }
return c
}
function col_build() -> void {
col_side = (TERRAIN_HALF * 2) / COL_CELL
let ncell = col_side * col_side
if col_start == null { col_start = words(ncell + 1); col_sorted = words(COL_CAP) }
for i in 0 .. ncell + 1 { col_start[i] = 0 }
for i in 0 .. col_n { col_start[col_cell_of(col_z[i], ter_oz) * col_side + col_cell_of(col_x[i], ter_ox) + 1] += 1 }
for c in 0 .. ncell { col_start[c + 1] += col_start[c] }
let fill = words(ncell)
for c in 0 .. ncell { fill[c] = col_start[c] }
for i in 0 .. col_n {
let c = col_cell_of(col_z[i], ter_oz) * col_side + col_cell_of(col_x[i], ter_ox)
col_sorted[fill[c]] = i
fill[c] += 1
}
free(fill)
col_built = true
print(`colliders: {col_n}`)
}
# push (px, pz) with radius pr out of every circle it overlaps; the result is in col_out
function col_resolve(px: int, pz: int, pr: int) -> bool {
col_out[0] = px; col_out[1] = pz
if not col_built or col_n == 0 { return false }
var x = px; var z = pz
var moved = false
let cx = col_cell_of(px, ter_ox); let cz = col_cell_of(pz, ter_oz)
for pass in 0 .. 2 {
for dz in 0 .. 3 {
let zc = cz + dz - 1
if zc < 0 or zc >= col_side { continue }
for dx in 0 .. 3 {
let xc = cx + dx - 1
if xc < 0 or xc >= col_side { continue }
let c = zc * col_side + xc
for k in col_start[c] .. col_start[c + 1] {
let i = col_sorted[k]
let ex = f_sub(x, col_x[i]); let ez = f_sub(z, col_z[i])
let d2 = f_add(f_mul(ex, ex), f_mul(ez, ez))
let rr = f_add(col_r[i], pr)
if f_ls(d2, f_mul(rr, rr)) {
var d = f_sqrt(d2)
var nx = ex; var nz = ez
if f_ls(d, fl(0.001)) { d = fl(0.001); nx = F_ONE; nz = F_ZERO }
let push = f_sub(rr, d)
x = f_add(x, f_mul(f_div(nx, d), push))
z = f_add(z, f_mul(f_div(nz, d), push))
moved = true
}
}
}
}
}
col_out[0] = x; col_out[1] = z
return moved
}
# is the segment from (x0,z0) to (x1,z1) clear of every circle (a camera line of sight)?
function col_clear(x0: int, z0: int, x1: int, z1: int, r: int) -> bool {
let steps = 6
for s in 0 .. steps + 1 {
let t = fr(s, steps)
let x = f_lerp(x0, x1, t); let z = f_lerp(z0, z1, t)
if col_resolve(x, z, r) { return false }
}
return true
}

View file

@ -0,0 +1,162 @@
# ============================================================================
# daylight.ludic — a time of day over the HDRI sky. The photograph is one late
# morning; the game wants a whole day. The sun's direction and colour follow a
# simple solar arc from the time (`daylight_set`), the sky image turns to keep
# its disc under that sun, the image-based light is scaled toward a deep-blue
# night (`u_ibl_scale`), and after dusk the light becomes a moon so the world
# still has shadows and form. A campfire is the one point light (`u_fire_*`).
# The convolved sky maps are rebuilt only when the sky has turned far from the
# angle they were baked at; the height-field shadow rebakes as the sun moves.
# ============================================================================
var day_on: bool = false
var day_hours: int = 0 # float bits, 0 .. 24
var day_light: int = 0 # 0 night .. 1 full day (float bits)
var day_ibl: words = null # rgb scale on the sky's light
var day_sun_base: words = null # the HDRI's sun radiance, kept
var day_az0: int = 0 # the sun's azimuth at the reference hour (radians, yaw convention)
var day_yaw0: int = 0 # the sky yaw the scene was tuned at
var day_hour0: int = 0 # the hour the photograph was taken (10.5)
var day_dir: int = 0 # +1 / -1: which way the sun travels in yaw
var day_az: int = 0
var day_el: int = 0
var day_gen: int = 0 # bumps when the light moved enough to rebake the terrain shadow
var day_baked_az: int = 0
var day_baked_el: int = 0
var day_sky_baked: int = 0 # the sky yaw the convolutions were baked at
var fire_pos: words = null
var fire_color: words = null
var day_moon: bool = false
var hand_pos: words = null
var hand_color: words = null
var hand_dir: words = null
var hand_cone: int = 0
var day_overcast: int = 0 # 0 clear .. 1 a low grey sky (float bits): dims the sun and the sky's light
var day_flash: int = 0 # a lightning flash this frame (0..1): the sun brightens for it
var day_fog_mul: int = 0x3F800000 # multiplies the base fog density (rain and snow thicken the air)
var day_fog_base: int = 0
function daylight_init() -> void {
day_ibl = v3_new(F_ONE, F_ONE, F_ONE)
day_sun_base = v3_new(sun_color[0], sun_color[1], sun_color[2])
fire_pos = v3_new(F_ZERO, fi(-1000), F_ZERO)
fire_color = v3_new(F_ZERO, F_ZERO, F_ZERO)
hand_pos = v3_new(F_ZERO, fi(-1000), F_ZERO); hand_color = v3_new(F_ZERO, F_ZERO, F_ZERO); hand_dir = v3_new(F_ZERO, F_ZERO, f_neg(F_ONE)); hand_cone = f_neg(F_TWO)
day_light = F_ONE
}
# Start the clock: the scene as tuned (sky yaw `yaw0`) is the photograph's hour `hour0`;
# the sun rises and sets toward `set_yaw` (the direction it should be in at 19:00).
function daylight_start(yaw0: int, hour0: int, set_yaw: int) -> void {
day_yaw0 = yaw0; day_hour0 = hour0
day_az0 = f_atan2(f_neg(sun_dir[0]), f_neg(sun_dir[2]))
day_sky_baked = yaw0
# which way does the sun travel? the way that puts it nearest `set_yaw` at 19:00
let step = f_mul(f_sub(fl(19.0), hour0), f_rad(fi(15)))
let da = f_abs(day_wrap(f_sub(f_add(day_az0, step), set_yaw)))
let db = f_abs(day_wrap(f_sub(f_sub(day_az0, step), set_yaw)))
day_dir = F_ONE
if f_ls(db, da) { day_dir = f_neg(F_ONE) }
day_on = true
day_baked_az = fi(1000)
daylight_set(hour0)
}
function day_wrap(a: int) -> int {
let two_pi = f_mul(F_TWO, F_PI)
var d = a
while f_gt(d, F_PI) { d = f_sub(d, two_pi) }
while f_ls(d, f_neg(F_PI)) { d = f_add(d, two_pi) }
return d
}
function smoothf(a: int, b: int, x: int) -> int {
let t = f_clamp(f_div(f_sub(x, a), f_sub(b, a)), F_ZERO, F_ONE)
return f_mul(f_mul(t, t), f_sub(fi(3), f_mul(F_TWO, t)))
}
function daylight_set(hours: int) -> void {
var h = f_mod(hours, fi(24))
if f_ls(h, F_ZERO) { h = f_add(h, fi(24)) }
day_hours = h
# the arc: up at 5:30, highest (about 57 degrees) at 12:45, down at 20:00
let t = f_mul(f_div(f_sub(h, fl(5.5)), fl(14.5)), F_PI)
let el = f_rad(f_add(fi(-6), f_mul(fi(63), f_sin(t))))
let az = f_add(day_az0, f_mul(f_mul(f_sub(h, day_hour0), f_rad(fi(15))), day_dir))
day_az = az; day_el = el
let d = smoothf(f_rad(fi(-8)), f_rad(fi(12)), el)
day_light = d
# the sun, warm and dim near the horizon; past dusk, the moon from across the sky
var laz = az; var lel = f_max(el, f_rad(fi(3)))
var warm_r = F_ONE; var warm_g = F_ONE; var warm_b = F_ONE
let low = smoothf(F_ZERO, f_rad(fi(24)), el)
warm_g = f_lerp(fl(0.55), F_ONE, low); warm_b = f_lerp(fl(0.28), F_ONE, low)
# an overcast sky: the sun goes diffuse and grey, a lightning flash brings it back white
let oc = f_clamp(day_overcast, F_ZERO, F_ONE)
let sunk = f_add(f_sub(F_ONE, f_mul(fl(0.92), oc)), f_mul(fl(2.5), day_flash))
var sr = f_mul(day_sun_base[0], f_mul(f_mul(d, warm_r), sunk))
var sg = f_mul(day_sun_base[1], f_mul(f_mul(d, warm_g), sunk))
var sb = f_mul(day_sun_base[2], f_mul(f_mul(d, warm_b), sunk))
day_moon = false
if f_ls(el, f_rad(fi(-7))) {
day_moon = true
laz = f_add(az, F_PI)
lel = f_clamp(f_neg(el), f_rad(fi(10)), f_rad(fi(55)))
let m = smoothf(f_rad(fi(-7)), f_rad(fi(-16)), el)
sr = f_mul(day_sun_base[0], f_mul(fl(0.0045), m))
sg = f_mul(day_sun_base[1], f_mul(fl(0.0060), m))
sb = f_mul(day_sun_base[2], f_mul(fl(0.0095), m))
}
let ce = f_cos(lel)
v3_set(sun_dir, f_neg(f_mul(f_sin(laz), ce)), f_sin(lel), f_neg(f_mul(f_cos(laz), ce)))
v3_set(sun_color, sr, sg, sb)
# the sky's light: full by day, a deep blue by night, amber through the dusk
let dusk = f_mul(smoothf(f_rad(fi(-10)), f_rad(fi(2)), el), f_sub(F_ONE, smoothf(f_rad(fi(2)), f_rad(fi(18)), el)))
v3_set(day_ibl, f_lerp(fl(0.020), F_ONE, d), f_lerp(fl(0.026), F_ONE, d), f_lerp(fl(0.045), F_ONE, d))
day_ibl[0] = f_mul(day_ibl[0], f_add(F_ONE, f_mul(fl(0.35), dusk)))
day_ibl[2] = f_mul(day_ibl[2], f_sub(F_ONE, f_mul(fl(0.25), dusk)))
# clouds: less light, and greyer (the blue and the warmth both fade)
let grey = f_add(f_mul(f_add(day_ibl[0], f_add(day_ibl[1], day_ibl[2])), fl(0.3333)), f_mul(fl(0.6), day_flash))
let dim = f_sub(F_ONE, f_mul(fl(0.55), oc))
for i in 0 .. 3 { day_ibl[i] = f_mul(f_lerp(day_ibl[i], grey, f_mul(fl(0.7), oc)), dim) }
if day_fog_base == 0 { day_fog_base = r3d_fog_density }
r3d_fog_density = f_mul(day_fog_base, day_fog_mul)
# exposure: auto-exposure must not turn the night into day
post_exposure_max = f_lerp(fl(4.5), fi(20), d)
# the visible sky turns with the sun (cheap); its convolutions rebake when far off
let sky_yaw_now = f_add(day_yaw0, f_mul(f_mul(f_sub(h, day_hour0), f_rad(fi(15))), day_dir))
sky_set_rot(sky_yaw_now)
if f_gt(f_abs(day_wrap(f_sub(sky_yaw_now, day_sky_baked))), f_rad(fi(35))) and f_gt(d, fl(0.05)) {
day_sky_baked = sky_yaw_now
sky_precompute()
}
# the terrain's baked shadow follows the light in steps
if f_gt(f_abs(day_wrap(f_sub(laz, day_baked_az))), f_rad(fi(4))) or f_gt(f_abs(f_sub(lel, day_baked_el)), f_rad(fi(3))) {
day_baked_az = laz; day_baked_el = lel
day_gen += 1
}
}
# the weather over the valley: overcast 0..1, a fog multiplier, a lightning flash 0..1
function daylight_weather(overcast: int, fog_mul: int, flash: int) -> void {
day_overcast = overcast; day_fog_mul = fog_mul; day_flash = flash
}
# the campfire: a point light at (x, y, z) of `strength` (0 = out)
function daylight_fire(x: int, y: int, z: int, strength: int) -> void {
v3_set(fire_pos, x, y, z)
v3_set(fire_color, f_mul(fl(9.0), strength), f_mul(fl(4.6), strength), f_mul(fl(1.4), strength))
}
# the light in the hand: a point light (cone < -1) or a cone along dir (cone = cos half-angle)
function daylight_hand(x: int, y: int, z: int, dx: int, dy: int, dz: int, cone: int, r: int, g: int, b: int) -> void {
v3_set(hand_pos, x, y, z); v3_set(hand_dir, dx, dy, dz); hand_cone = cone
v3_set(hand_color, r, g, b)
}
function daylight_bind(prog: int) -> void {
u_v3(gl_uniform(prog, "u_hand_pos"), hand_pos)
u_v3(gl_uniform(prog, "u_hand_color"), hand_color)
u_v3(gl_uniform(prog, "u_hand_dir"), hand_dir)
u_f(gl_uniform(prog, "u_hand_cone"), hand_cone)
u_v3(gl_uniform(prog, "u_ibl_scale"), day_ibl)
u_f(gl_uniform(prog, "u_daylight"), day_light)
u_v3(gl_uniform(prog, "u_fire_pos"), fire_pos)
u_v3(gl_uniform(prog, "u_fire_color"), fire_color)
}

View file

@ -0,0 +1,206 @@
# ============================================================================
# fmath.ludic — IEEE single-precision math for the renderer.
#
# Ludic's own numbers are Q16.16; a renderer wants the floats the GPU eats. A
# float lives here as its bit pattern in an `int`, arithmetic goes through the
# f_* helpers Gl.* provides (gl.ll), and vectors / matrices are `words` buffers
# of those bit patterns — which is exactly the memory layout glUniform* and
# glBufferData expect, so nothing is converted at upload.
#
# fl(x) a fixed literal as a float fr(n, d) n/d as a float
# v3_* 3-vectors at an index of a words buffer (x, y, z consecutive)
# m4_* 4x4 column-major matrices in a 16-word buffer
# ============================================================================
const F_ZERO: int = 0x00000000
const F_ONE: int = 0x3F800000
const F_TWO: int = 0x40000000
const F_HALF: int = 0x3F000000
const F_PI: int = 0x40490FDB
function fl(x: fixed) -> int { return fx_to_f32(x) }
function fi(n: int) -> int { return f_from_int(n) }
function fr(n: int, d: int) -> int { return f_div(f_from_int(n), f_from_int(d)) }
function f_neg1() -> int { return f_neg(F_ONE) }
function f_clamp(x: int, lo: int, hi: int) -> int { return f_min(f_max(x, lo), hi) }
function f_lerp(a: int, b: int, t: int) -> int { return f_add(a, f_mul(f_sub(b, a), t)) }
function f_gt(a: int, b: int) -> bool { return f_lt(b, a) != 0 }
function f_ls(a: int, b: int) -> bool { return f_lt(a, b) != 0 }
function f_rad(deg: int) -> int { return f_mul(deg, f_div(F_PI, fi(180))) }
function f_fx(x: int) -> fixed { return f32_to_fx(x) }
# ---- vectors ----------------------------------------------------------------
function v3_new(x: int, y: int, z: int) -> words {
let v = words(3)
v[0] = x
v[1] = y
v[2] = z
return v
}
function v3_set(v: words, x: int, y: int, z: int) -> void { v[0] = x; v[1] = y; v[2] = z }
function v3_copy(o: words, a: words) -> void { o[0] = a[0]; o[1] = a[1]; o[2] = a[2] }
function v3_add(o: words, a: words, b: words) -> void {
o[0] = f_add(a[0], b[0]); o[1] = f_add(a[1], b[1]); o[2] = f_add(a[2], b[2])
}
function v3_sub(o: words, a: words, b: words) -> void {
o[0] = f_sub(a[0], b[0]); o[1] = f_sub(a[1], b[1]); o[2] = f_sub(a[2], b[2])
}
function v3_scale(o: words, a: words, s: int) -> void {
o[0] = f_mul(a[0], s); o[1] = f_mul(a[1], s); o[2] = f_mul(a[2], s)
}
function v3_madd(o: words, a: words, b: words, s: int) -> void { # o = a + b*s
o[0] = f_add(a[0], f_mul(b[0], s)); o[1] = f_add(a[1], f_mul(b[1], s)); o[2] = f_add(a[2], f_mul(b[2], s))
}
function v3_dot(a: words, b: words) -> int {
return f_add(f_add(f_mul(a[0], b[0]), f_mul(a[1], b[1])), f_mul(a[2], b[2]))
}
function v3_cross(o: words, a: words, b: words) -> void {
let x = f_sub(f_mul(a[1], b[2]), f_mul(a[2], b[1]))
let y = f_sub(f_mul(a[2], b[0]), f_mul(a[0], b[2]))
let z = f_sub(f_mul(a[0], b[1]), f_mul(a[1], b[0]))
o[0] = x; o[1] = y; o[2] = z
}
function v3_len(a: words) -> int { return f_sqrt(v3_dot(a, a)) }
function v3_normalize(o: words, a: words) -> void {
let l = v3_len(a)
if l == 0 { o[0] = 0; o[1] = 0; o[2] = 0; return }
let inv = f_div(F_ONE, l)
v3_scale(o, a, inv)
}
function v3_dist(a: words, b: words) -> int {
let t = words(3)
v3_sub(t, a, b)
let d = v3_len(t)
free(t)
return d
}
# ---- matrices (column-major, m[col*4 + row]) ------------------------------------
function m4_new() -> words { let m = words(16); m4_identity(m); return m }
function m4_identity(m: words) -> void {
for i in 0 .. 16 { m[i] = F_ZERO }
m[0] = F_ONE; m[5] = F_ONE; m[10] = F_ONE; m[15] = F_ONE
}
function m4_copy(o: words, a: words) -> void { for i in 0 .. 16 { o[i] = a[i] } }
# o = a * b (o may not alias a or b)
function m4_mul(o: words, a: words, b: words) -> void {
for c in 0 .. 4 {
for r in 0 .. 4 {
var s = F_ZERO
for k in 0 .. 4 { s = f_add(s, f_mul(a[k * 4 + r], b[c * 4 + k])) }
o[c * 4 + r] = s
}
}
}
function m4_mul_into(a: words, b: words) -> void { # a = a * b
let t = words(16)
m4_mul(t, a, b)
m4_copy(a, t)
free(t)
}
function m4_translation(m: words, x: int, y: int, z: int) -> void {
m4_identity(m)
m[12] = x; m[13] = y; m[14] = z
}
function m4_scaling(m: words, x: int, y: int, z: int) -> void {
m4_identity(m)
m[0] = x; m[5] = y; m[10] = z
}
function m4_rotation_y(m: words, angle: int) -> void {
m4_identity(m)
let c = f_cos(angle); let s = f_sin(angle)
m[0] = c; m[2] = f_neg(s); m[8] = s; m[10] = c
}
function m4_rotation_x(m: words, angle: int) -> void {
m4_identity(m)
let c = f_cos(angle); let s = f_sin(angle)
m[5] = c; m[6] = s; m[9] = f_neg(s); m[10] = c
}
function m4_rotation_z(m: words, angle: int) -> void {
m4_identity(m)
let c = f_cos(angle); let s = f_sin(angle)
m[0] = c; m[1] = s; m[4] = f_neg(s); m[5] = c
}
# a model matrix: translate * rotate_y * uniform scale
function m4_trs(m: words, x: int, y: int, z: int, yaw: int, s: int) -> void {
m4_rotation_y(m, yaw)
for i in 0 .. 12 { m[i] = f_mul(m[i], s) }
m[12] = x; m[13] = y; m[14] = z
}
# OpenGL clip space (z in [-1, 1]); fovy in radians
function m4_perspective(m: words, fovy: int, aspect: int, near: int, far: int) -> void {
for i in 0 .. 16 { m[i] = F_ZERO }
let f = f_div(F_ONE, f_tan(f_mul(fovy, F_HALF)))
m[0] = f_div(f, aspect)
m[5] = f
m[10] = f_div(f_add(far, near), f_sub(near, far))
m[11] = f_neg(F_ONE)
m[14] = f_div(f_mul(f_mul(F_TWO, far), near), f_sub(near, far))
}
function m4_ortho(m: words, l: int, r: int, b: int, t: int, n: int, f: int) -> void {
m4_identity(m)
m[0] = f_div(F_TWO, f_sub(r, l))
m[5] = f_div(F_TWO, f_sub(t, b))
m[10] = f_div(f_neg(F_TWO), f_sub(f, n))
m[12] = f_neg(f_div(f_add(r, l), f_sub(r, l)))
m[13] = f_neg(f_div(f_add(t, b), f_sub(t, b)))
m[14] = f_neg(f_div(f_add(f, n), f_sub(f, n)))
}
function m4_look_at(m: words, eye: words, at: words, up: words) -> void {
let fwd = words(3); let side = words(3); let u = words(3); let t = words(3)
v3_sub(t, at, eye)
v3_normalize(fwd, t)
v3_cross(t, fwd, up)
v3_normalize(side, t)
v3_cross(u, side, fwd)
m4_identity(m)
m[0] = side[0]; m[4] = side[1]; m[8] = side[2]
m[1] = u[0]; m[5] = u[1]; m[9] = u[2]
m[2] = f_neg(fwd[0]); m[6] = f_neg(fwd[1]); m[10] = f_neg(fwd[2])
m[12] = f_neg(v3_dot(side, eye))
m[13] = f_neg(v3_dot(u, eye))
m[14] = v3_dot(fwd, eye)
free(fwd); free(side); free(u); free(t)
}
# general 4x4 inverse (cofactor expansion); o may not alias a
function m4_inverse(o: words, a: words) -> bool {
let inv = words(16)
inv[0] = f_add(f_sub(f_add(f_mul(a[5], f_mul(a[10], a[15])), f_neg(f_mul(a[5], f_mul(a[11], a[14])))), f_mul(a[9], f_mul(a[6], a[15]))), f_add(f_mul(a[9], f_mul(a[7], a[14])), f_sub(f_mul(a[13], f_mul(a[6], a[11])), f_mul(a[13], f_mul(a[7], a[10])))))
inv[4] = f_add(f_sub(f_add(f_neg(f_mul(a[4], f_mul(a[10], a[15]))), f_mul(a[4], f_mul(a[11], a[14]))), f_mul(a[8], f_mul(a[7], a[14]))), f_add(f_mul(a[8], f_mul(a[6], a[15])), f_sub(f_mul(a[12], f_mul(a[7], a[10])), f_mul(a[12], f_mul(a[6], a[11])))))
inv[8] = f_add(f_sub(f_add(f_mul(a[4], f_mul(a[9], a[15])), f_neg(f_mul(a[4], f_mul(a[11], a[13])))), f_mul(a[8], f_mul(a[5], a[15]))), f_add(f_mul(a[8], f_mul(a[7], a[13])), f_sub(f_mul(a[12], f_mul(a[5], a[11])), f_mul(a[12], f_mul(a[7], a[9])))))
inv[12] = f_add(f_sub(f_add(f_neg(f_mul(a[4], f_mul(a[9], a[14]))), f_mul(a[4], f_mul(a[10], a[13]))), f_mul(a[8], f_mul(a[6], a[13]))), f_add(f_mul(a[8], f_mul(a[5], a[14])), f_sub(f_mul(a[12], f_mul(a[6], a[9])), f_mul(a[12], f_mul(a[5], a[10])))))
inv[1] = f_add(f_sub(f_add(f_neg(f_mul(a[1], f_mul(a[10], a[15]))), f_mul(a[1], f_mul(a[11], a[14]))), f_mul(a[9], f_mul(a[3], a[14]))), f_add(f_mul(a[9], f_mul(a[2], a[15])), f_sub(f_mul(a[13], f_mul(a[3], a[10])), f_mul(a[13], f_mul(a[2], a[11])))))
inv[5] = f_add(f_sub(f_add(f_mul(a[0], f_mul(a[10], a[15])), f_neg(f_mul(a[0], f_mul(a[11], a[14])))), f_mul(a[8], f_mul(a[2], a[15]))), f_add(f_mul(a[8], f_mul(a[3], a[14])), f_sub(f_mul(a[12], f_mul(a[2], a[11])), f_mul(a[12], f_mul(a[3], a[10])))))
inv[9] = f_add(f_sub(f_add(f_neg(f_mul(a[0], f_mul(a[9], a[15]))), f_mul(a[0], f_mul(a[11], a[13]))), f_mul(a[8], f_mul(a[3], a[13]))), f_add(f_mul(a[8], f_mul(a[1], a[15])), f_sub(f_mul(a[12], f_mul(a[3], a[9])), f_mul(a[12], f_mul(a[1], a[11])))))
inv[13] = f_add(f_sub(f_add(f_mul(a[0], f_mul(a[9], a[14])), f_neg(f_mul(a[0], f_mul(a[10], a[13])))), f_mul(a[8], f_mul(a[1], a[14]))), f_add(f_mul(a[8], f_mul(a[2], a[13])), f_sub(f_mul(a[12], f_mul(a[1], a[10])), f_mul(a[12], f_mul(a[2], a[9])))))
inv[2] = f_add(f_sub(f_add(f_mul(a[1], f_mul(a[6], a[15])), f_neg(f_mul(a[1], f_mul(a[7], a[14])))), f_mul(a[5], f_mul(a[2], a[15]))), f_add(f_mul(a[5], f_mul(a[3], a[14])), f_sub(f_mul(a[13], f_mul(a[2], a[7])), f_mul(a[13], f_mul(a[3], a[6])))))
inv[6] = f_add(f_sub(f_add(f_neg(f_mul(a[0], f_mul(a[6], a[15]))), f_mul(a[0], f_mul(a[7], a[14]))), f_mul(a[4], f_mul(a[3], a[14]))), f_add(f_mul(a[4], f_mul(a[2], a[15])), f_sub(f_mul(a[12], f_mul(a[3], a[6])), f_mul(a[12], f_mul(a[2], a[7])))))
inv[10] = f_add(f_sub(f_add(f_mul(a[0], f_mul(a[5], a[15])), f_neg(f_mul(a[0], f_mul(a[7], a[13])))), f_mul(a[4], f_mul(a[1], a[15]))), f_add(f_mul(a[4], f_mul(a[3], a[13])), f_sub(f_mul(a[12], f_mul(a[1], a[7])), f_mul(a[12], f_mul(a[3], a[5])))))
inv[14] = f_add(f_sub(f_add(f_neg(f_mul(a[0], f_mul(a[5], a[14]))), f_mul(a[0], f_mul(a[6], a[13]))), f_mul(a[4], f_mul(a[2], a[13]))), f_add(f_mul(a[4], f_mul(a[1], a[14])), f_sub(f_mul(a[12], f_mul(a[2], a[5])), f_mul(a[12], f_mul(a[1], a[6])))))
inv[3] = f_add(f_sub(f_add(f_neg(f_mul(a[1], f_mul(a[6], a[11]))), f_mul(a[1], f_mul(a[7], a[10]))), f_mul(a[5], f_mul(a[3], a[10]))), f_add(f_mul(a[5], f_mul(a[2], a[11])), f_sub(f_mul(a[9], f_mul(a[3], a[6])), f_mul(a[9], f_mul(a[2], a[7])))))
inv[7] = f_add(f_sub(f_add(f_mul(a[0], f_mul(a[6], a[11])), f_neg(f_mul(a[0], f_mul(a[7], a[10])))), f_mul(a[4], f_mul(a[2], a[11]))), f_add(f_mul(a[4], f_mul(a[3], a[10])), f_sub(f_mul(a[8], f_mul(a[2], a[7])), f_mul(a[8], f_mul(a[3], a[6])))))
inv[11] = f_add(f_sub(f_add(f_neg(f_mul(a[0], f_mul(a[5], a[11]))), f_mul(a[0], f_mul(a[7], a[9]))), f_mul(a[4], f_mul(a[3], a[9]))), f_add(f_mul(a[4], f_mul(a[1], a[11])), f_sub(f_mul(a[8], f_mul(a[3], a[5])), f_mul(a[8], f_mul(a[1], a[7])))))
inv[15] = f_add(f_sub(f_add(f_mul(a[0], f_mul(a[5], a[10])), f_neg(f_mul(a[0], f_mul(a[6], a[9])))), f_mul(a[4], f_mul(a[1], a[10]))), f_add(f_mul(a[4], f_mul(a[2], a[9])), f_sub(f_mul(a[8], f_mul(a[1], a[6])), f_mul(a[8], f_mul(a[2], a[5])))))
let det = f_add(f_add(f_mul(a[0], inv[0]), f_mul(a[1], inv[4])), f_add(f_mul(a[2], inv[8]), f_mul(a[3], inv[12])))
if det == 0 { free(inv); return false }
let id = f_div(F_ONE, det)
for i in 0 .. 16 { o[i] = f_mul(inv[i], id) }
free(inv)
return true
}
# transform a point (w = 1) by m: o = m * (x, y, z, 1), returns w
function m4_xform_point(o: words, m: words, x: int, y: int, z: int) -> int {
o[0] = f_add(f_add(f_mul(m[0], x), f_mul(m[4], y)), f_add(f_mul(m[8], z), m[12]))
o[1] = f_add(f_add(f_mul(m[1], x), f_mul(m[5], y)), f_add(f_mul(m[9], z), m[13]))
o[2] = f_add(f_add(f_mul(m[2], x), f_mul(m[6], y)), f_add(f_mul(m[10], z), m[14]))
return f_add(f_add(f_mul(m[3], x), f_mul(m[7], y)), f_add(f_mul(m[11], z), m[15]))
}
# ---- uniforms ------------------------------------------------------------------
function u_mat4(loc: int, m: words) -> void { gl_uniform_matrix4fv(loc, 1, 0, m) }
function u_f(loc: int, v: int) -> void { let t = gl_scratch(); t[0] = v; gl_uniform1fv(loc, 1, t) }
function u_f2(loc: int, x: int, y: int) -> void { let t = gl_scratch(); t[0] = x; t[1] = y; gl_uniform2fv(loc, 1, t) }
function u_f3(loc: int, x: int, y: int, z: int) -> void { let t = gl_scratch(); t[0] = x; t[1] = y; t[2] = z; gl_uniform3fv(loc, 1, t) }
function u_f4(loc: int, x: int, y: int, z: int, w: int) -> void { let t = gl_scratch(); t[0] = x; t[1] = y; t[2] = z; t[3] = w; gl_uniform4fv(loc, 1, t) }
function u_v3(loc: int, v: words) -> void { gl_uniform3fv(loc, 1, v) }
function u_i(loc: int, v: int) -> void { gl_uniform1i(loc, v) }

View file

@ -0,0 +1,229 @@
# ============================================================================
# gltf.ludic — a glTF 2.0 loader for scanned models: one named node's mesh,
# each primitive uploaded straight from the .bin (positions, normals, uvs,
# indices) with its baseColor / normal / ARM textures. Only the byte ranges the
# node needs are read, so a 500 MB scan library costs what one tree costs.
# ============================================================================
property Prim {
mesh: Mesh,
diff: int = 0,
nrm: int = 0,
arm: int = 0,
name: string # the material's name (a part to tint: "hk_jacket")
}
property Model {
prims: []Prim,
radius: int = 0, # float bits: max horizontal extent from the origin
height: int = 0, # float bits: y extent above ymin
ymin: int = 0,
tris: int = 0,
skin: Skin # the skeleton, for a skinned node (skin.ludic); null for a rigid model
}
var gltf_dir: string = null
var gltf_bin: pointer = null
var gltf_doc: Val = null
var gltf_count: int = 0
var gltf_ctype: int = 0
var gltf_comps: int = 0
var gltf_tex_paths: []pointer = null
var gltf_tex_ids: words = null
var gltf_tex_n: int = 0
var gltf_white: int = 0
var gltf_flat: int = 0
# a JSON number as float bits (ints and fixed-point decimals both)
function jnum(v: Val) -> int {
if v.tag == 2 { return fx_to_f32(v.num) }
return f_from_int(v.num)
}
function jint(v: Val, key: pointer, fallback: int) -> int {
if value_has(v, key) == 0 { return fallback }
return value_as_int(value_get(v, key))
}
# Materials seen so far, by name: a LOD chain exported from the .blend (tools/glgen/
# lod_export.py) carries the material names but no images — the blend links textures
# that are not downloaded with it — so its primitives take the textures the LOD0
# download's material of the same name loaded.
var gltf_mat_names: []string = null
var gltf_mat_diff: words = null
var gltf_mat_nrm: words = null
var gltf_mat_arm: words = null
var gltf_mat_n: int = 0
function gltf_mat_find(name: string) -> int {
if gltf_mat_names == null { return -1 }
for i in 0 .. gltf_mat_n { if gltf_mat_names[i] == name { return i } }
return -1
}
function gltf_mat_remember(name: string, diff: int, nrm: int, arm: int) -> void {
if gltf_mat_names == null { gltf_mat_names = new []string; gltf_mat_diff = words(256); gltf_mat_nrm = words(256); gltf_mat_arm = words(256) }
if gltf_mat_find(name) >= 0 or gltf_mat_n >= 256 { return }
push(gltf_mat_names, name); gltf_mat_diff[gltf_mat_n] = diff; gltf_mat_nrm[gltf_mat_n] = nrm; gltf_mat_arm[gltf_mat_n] = arm
gltf_mat_n += 1
}
var gltf_cutout: bool = false # the material being loaded is alpha-blended (a cut-out atlas)
function gltf_texture(uri: pointer, srgb: bool) -> int {
if gltf_tex_paths == null { gltf_tex_paths = new []pointer; gltf_tex_ids = words(256) }
var png: string = uri
let n = len(uri)
if n > 4 and uri[n - 4] == '.' and uri[n - 3] == 'j' { png = uri[0 .. n - 4] + ".png" }
let path = gltf_dir + "/" + png
var i = 0
while i < gltf_tex_n { if gltf_tex_paths[i] == path { return gltf_tex_ids[i] }; i += 1 }
var dil = 0
if gltf_cutout { dil = 24 }
let id = tex_load_ex(path, srgb, dil)
if gltf_tex_n < 256 { push(gltf_tex_paths, path); gltf_tex_ids[gltf_tex_n] = id; gltf_tex_n += 1 }
return id
}
# the image uri behind materials[m].<slot>.index, or null
function gltf_mat_uri(mat: Val, slot: pointer) -> pointer {
var holder = mat
if slot == "baseColorTexture" or slot == "metallicRoughnessTexture" {
if value_has(mat, "pbrMetallicRoughness") == 0 { return null }
holder = value_get(mat, "pbrMetallicRoughness")
}
if value_has(holder, slot) == 0 { return null }
let ti = value_as_int(value_get(value_get(holder, slot), "index"))
let tex = value_at(value_get(gltf_doc, "textures"), ti)
let src = value_as_int(value_get(tex, "source"))
let img = value_at(value_get(gltf_doc, "images"), src)
return value_as_str(value_get(img, "uri"))
}
# read one accessor's bytes from the .bin; sets gltf_count / gltf_ctype / gltf_comps
function gltf_accessor(idx: int) -> pointer {
let acc = value_at(value_get(gltf_doc, "accessors"), idx)
let bv = value_at(value_get(gltf_doc, "bufferViews"), jint(acc, "bufferView", 0))
let off = jint(bv, "byteOffset", 0) + jint(acc, "byteOffset", 0)
gltf_count = jint(acc, "count", 0)
gltf_ctype = jint(acc, "componentType", 5126)
let ty = value_as_str(value_get(acc, "type"))
gltf_comps = 1
if ty == "VEC2" { gltf_comps = 2 }
if ty == "VEC3" { gltf_comps = 3 }
if ty == "VEC4" { gltf_comps = 4 }
if ty == "MAT2" { gltf_comps = 4 }
if ty == "MAT3" { gltf_comps = 9 }
if ty == "MAT4" { gltf_comps = 16 } # a skin's inverse bind matrices
var csz = 4
if gltf_ctype == 5123 or gltf_ctype == 5122 { csz = 2 }
if gltf_ctype == 5121 or gltf_ctype == 5120 { csz = 1 }
let n = gltf_count * gltf_comps * csz
let buf = bytes(n + 8)
file_seek(gltf_bin, off, 0)
file_read(gltf_bin, buf, n)
return buf
}
function gltf_attrib(m: Mesh, attrs: Val, name: pointer, loc: int) -> bool {
if value_has(attrs, name) == 0 { return false }
let data = gltf_accessor(value_as_int(value_get(attrs, name)))
let b = gl_buffer()
gl_bind_buffer(GL_ARRAY_BUFFER, b)
gl_buffer_data(GL_ARRAY_BUFFER, gltf_count * gltf_comps * 4, data, GL_STATIC_DRAW)
gl_enable_vertex_attrib_array(loc)
gl_vertex_attrib_pointer(loc, gltf_comps, GL_FLOAT, 0, 0, null)
free(data)
if loc == 0 { m.vbo = b }
return true
}
function gltf_prim(p: Val) -> Prim {
let pr = new Prim
let m = new Mesh
m.vao = gl_vao()
let attrs = value_get(p, "attributes")
gltf_attrib(m, attrs, "POSITION", 0)
gltf_attrib(m, attrs, "NORMAL", 1)
gltf_attrib(m, attrs, "TEXCOORD_0", 2)
skin_attribs(m, attrs) # JOINTS_0 / WEIGHTS_0 onto 5 / 6, when the mesh has them
let idx = gltf_accessor(value_as_int(value_get(p, "indices")))
var isz = 4
m.itype = GL_UNSIGNED_INT
if gltf_ctype == 5123 { isz = 2; m.itype = GL_UNSIGNED_SHORT }
m.ebo = gl_buffer()
gl_bind_buffer(GL_ELEMENT_ARRAY_BUFFER, m.ebo)
gl_buffer_data(GL_ELEMENT_ARRAY_BUFFER, gltf_count * isz, idx, GL_STATIC_DRAW)
free(idx)
m.count = gltf_count
gl_bind_vertex_array(0)
pr.mesh = m
# material textures
if gltf_white == 0 { gltf_white = tex_solid(200, 200, 200, 255); gltf_flat = tex_solid(128, 128, 255, 255) }
pr.diff = gltf_white; pr.nrm = gltf_flat; pr.arm = gltf_white
if value_has(p, "material") != 0 {
let mat = value_at(value_get(gltf_doc, "materials"), value_as_int(value_get(p, "material")))
gltf_cutout = value_has(mat, "alphaMode") != 0
let ud = gltf_mat_uri(mat, "baseColorTexture")
let un = gltf_mat_uri(mat, "normalTexture")
let ua = gltf_mat_uri(mat, "metallicRoughnessTexture")
if ud != null { pr.diff = gltf_texture(ud, true) }
if un != null { pr.nrm = gltf_texture(un, false) }
if ua != null { pr.arm = gltf_texture(ua, false) }
var mname: string = null
if value_has(mat, "name") != 0 { mname = value_as_str(value_get(mat, "name")) }
pr.name = mname
if mname != null {
if ud != null { gltf_mat_remember(mname, pr.diff, pr.nrm, pr.arm) }
else {
let k = gltf_mat_find(mname)
if k >= 0 { pr.diff = gltf_mat_diff[k]; pr.nrm = gltf_mat_nrm[k]; pr.arm = gltf_mat_arm[k] }
else { print(`gltf: material {mname} has no textures and none were loaded before it`) }
}
}
}
return pr
}
# Load the mesh of the node called `node_name` from dir/file.
function gltf_load(dir: string, file: string, node_name: string) -> Model {
gltf_dir = dir
let text = Fs.read_text(dir + "/" + file)
if text == null { print(`gltf: cannot read {dir}/{file}`); return null }
gltf_doc = Json.parse(text)
let buffers = value_get(gltf_doc, "buffers")
let bin_uri = value_as_str(value_get(value_at(buffers, 0), "uri"))
gltf_bin = file_open(dir + "/" + bin_uri, "rb")
if gltf_bin == null { print(`gltf: cannot open {bin_uri}`); return null }
let nodes = value_get(gltf_doc, "nodes")
var mesh_idx = -1
var skin_idx = -1
for i in 0 .. value_count(nodes) {
let nd = value_at(nodes, i)
if value_as_str(value_get(nd, "name")) == node_name { mesh_idx = jint(nd, "mesh", -1); skin_idx = jint(nd, "skin", -1) }
}
if mesh_idx < 0 { print(`gltf: no node {node_name} in {file}`); file_close(gltf_bin); return null }
let model = new Model
model.prims = new []Prim
let mesh = value_at(value_get(gltf_doc, "meshes"), mesh_idx)
let prims = value_get(mesh, "primitives")
var r2 = F_ZERO; var ymin = fi(1000); var ymax = fi(-1000)
for i in 0 .. value_count(prims) {
let p = value_at(prims, i)
push(model.prims, gltf_prim(p))
model.tris += gltf_count / 3
# bounds from the accessor min/max
let acc = value_at(value_get(gltf_doc, "accessors"), value_as_int(value_get(value_get(p, "attributes"), "POSITION")))
let mn = value_get(acc, "min"); let mx = value_get(acc, "max")
let x0 = f_abs(jnum(value_at(mn, 0))); let x1 = f_abs(jnum(value_at(mx, 0)))
let z0 = f_abs(jnum(value_at(mn, 2))); let z1 = f_abs(jnum(value_at(mx, 2)))
let rx = f_max(x0, x1); let rz = f_max(z0, z1)
let rr = f_add(f_mul(rx, rx), f_mul(rz, rz))
if f_gt(rr, r2) { r2 = rr }
let y0 = jnum(value_at(mn, 1)); let y1 = jnum(value_at(mx, 1))
if f_ls(y0, ymin) { ymin = y0 }
if f_gt(y1, ymax) { ymax = y1 }
}
if skin_idx >= 0 { model.skin = skin_load(skin_idx) }
file_close(gltf_bin)
model.radius = f_sqrt(r2)
model.ymin = ymin
model.height = f_sub(ymax, ymin)
print(`gltf: {node_name}: {len(model.prims)} prims, {model.tris} tris`)
return model
}

View file

@ -0,0 +1,161 @@
# ============================================================================
# grass.ludic — procedural GPU ground cover with continuous density (no rings).
#
# The world is cut into 16 m cells; blade j of a cell always stands in the same place
# (shaders/grass.vert). Draws are per tile: the CPU walks tiles around the camera,
# frustum-culls them, and feeds each visible tile as many blade indices per cell as its
# NEAREST point could need; the vertex stage then keeps only the indices that exist at
# each blade's own distance, so density is one smooth function of distance everywhere.
# Tiles are 16 m near, 64 m in the middle distance and 256 m far, purely to keep the
# draw count down — the cells and their hashes are the same in every tile size.
# ============================================================================
const GRASS_CELL: int = 16
var grass_prog: int = 0
var grass_mesh: Mesh = null
var grass_on: bool = true
var grass_wind: int = 0
var grass_s0: int = 0 # float bits: blade spacing at the camera (m)
var grass_d0: int = 0 # the distance at which the spacing has doubled (m)
var grass_radius: int = 0 # no blades past this (m)
var grass_draws: int = 0
var grass_dbg: int = 0
# a blade: `rows` rows of 2 vertices (x across, y along, z bend), attribute 2 = uv
function grass_blade_mesh(rows: int) -> Mesh {
let m = new Mesh
m.vao = gl_vao()
let v = gl_floats(rows * 2 * 5)
var k = 0
for r in 0 .. rows {
let t = fr(r, rows - 1)
let taper = f_max(f_sub(F_ONE, f_mul(t, f_mul(t, f_sqrt(t)))), fl(0.12))
let bend = f_mul(f_mul(t, t), fl(0.28))
for sd in 0 .. 2 {
var x = f_neg(F_HALF)
if sd == 1 { x = F_HALF }
gl_put_bits(v, k, f_mul(x, taper)); gl_put_bits(v, k + 1, t); gl_put_bits(v, k + 2, bend)
gl_put_bits(v, k + 3, fi(sd)); gl_put_bits(v, k + 4, t)
k += 5
}
}
m.vbo = gl_buffer()
gl_bind_buffer(GL_ARRAY_BUFFER, m.vbo)
gl_buffer_data(GL_ARRAY_BUFFER, gl_bytes_of(rows * 2 * 5), v, GL_STATIC_DRAW)
gl_enable_vertex_attrib_array(0); gl_vertex_attrib_pointer(0, 3, GL_FLOAT, 0, 20, null)
gl_enable_vertex_attrib_array(2); gl_vertex_attrib_pointer(2, 2, GL_FLOAT, 0, 20, gl_ptr(null, 12))
free(v)
let nq = rows - 1
let idx = words(nq * 6)
for q in 0 .. nq {
let b = q * 2
idx[q * 6] = b; idx[q * 6 + 1] = b + 1; idx[q * 6 + 2] = b + 2
idx[q * 6 + 3] = b + 1; idx[q * 6 + 4] = b + 3; idx[q * 6 + 5] = b + 2
}
m.ebo = gl_buffer()
gl_bind_buffer(GL_ELEMENT_ARRAY_BUFFER, m.ebo)
gl_buffer_data(GL_ELEMENT_ARRAY_BUFFER, nq * 6 * 4, idx, GL_STATIC_DRAW)
free(idx)
m.count = nq * 6
gl_bind_vertex_array(0)
return m
}
function grass_init() -> void {
grass_prog = r3d_program("grass.vert", "model.frag", "#define FOLIAGE\n#define BLADE\n")
grass_mesh = grass_blade_mesh(4)
grass_wind = fl(2.4)
grass_s0 = fl(0.11)
grass_d0 = fi(45)
grass_radius = fi(1600)
if Os.has_env("R3D_NOBLADES") { grass_on = false }
if Os.has_env("R3D_GRASS_R") { grass_radius = fi(Text.to_int(Os.env("R3D_GRASS_R"))) }
if Os.has_env("R3D_GRASS_DBG") { grass_dbg = Text.to_int(Os.env("R3D_GRASS_DBG")) }
}
# indices per 16 m cell that could exist at distance d (the count the shader computes)
function grass_count_at(d: int) -> int {
let spacing = f_mul(grass_s0, f_add(F_ONE, f_div(d, grass_d0)))
let n = f_div(fi(GRASS_CELL * GRASS_CELL), f_mul(spacing, spacing))
return f_to_int(n) + 1
}
# one tile size over one distance band
function grass_tiles(size: int, d_min: int, d_max: int) -> void {
let p = grass_prog
let cells = size / GRASS_CELL
gl_uniform1i(gl_uniform(p, "u_tile_cells"), cells)
let sz = fi(size)
let half = f_mul(sz, F_HALF)
let reach = f_add(d_max, f_mul(half, fl(1.5)))
let tx0 = f_to_int(f_floor(f_div(f_sub(cam_pos[0], reach), sz)))
let tx1 = f_to_int(f_floor(f_div(f_add(cam_pos[0], reach), sz)))
let tz0 = f_to_int(f_floor(f_div(f_sub(cam_pos[2], reach), sz)))
let tz1 = f_to_int(f_floor(f_div(f_add(cam_pos[2], reach), sz)))
let corner_r = f_mul(half, fl(1.42))
var tz = tz0
while tz <= tz1 {
var tx = tx0
while tx <= tx1 {
let ox = f_mul(fi(tx), sz); let oz = f_mul(fi(tz), sz)
let cx = f_add(ox, half); let cz = f_add(oz, half)
let dx = f_sub(cx, cam_pos[0]); let dz = f_sub(cz, cam_pos[2])
let dc = f_sqrt(f_add(f_mul(dx, dx), f_mul(dz, dz)))
# the tile's nearest and farthest points decide which band it belongs to
let dnear = f_max(f_sub(dc, corner_r), F_ZERO)
if f_ls(dc, d_min) or not f_ls(dnear, d_max) { tx += 1; continue }
let cy = terrain_height(cx, cz)
if cam_sphere_visible(cx, cy, cz, f_add(corner_r, fi(6))) {
let per = grass_count_at(dnear)
if per > 0 {
u_f2(gl_uniform(p, "u_tile"), ox, oz)
gl_uniform1i(gl_uniform(p, "u_per_cell"), per)
mesh_draw_instanced(grass_mesh, per * cells * cells)
grass_draws += 1
}
}
tx += 1
}
tz += 1
}
}
function grass_draw() -> void {
if not grass_on or ter_reflect or grass_prog == 0 { return }
let p = grass_prog
gl_use_program(p)
u_mat4(gl_uniform(p, "u_view"), cam_view)
u_mat4(gl_uniform(p, "u_proj"), cam_proj)
u_mat4(gl_uniform(p, "u_vp"), cam_vp_clean)
u_f(gl_uniform(p, "u_wind"), grass_wind)
u_f(gl_uniform(p, "u_rough_scale"), F_ONE)
u_v3(gl_uniform(p, "u_tint"), sc_blade_tint)
u_v3(gl_uniform(p, "u_blade_base"), sc_blade_base)
u_v3(gl_uniform(p, "u_blade_tip"), sc_blade_tip)
u_f(gl_uniform(p, "u_cull"), grass_radius)
u_f(gl_uniform(p, "u_model_h"), F_ZERO)
u_f(gl_uniform(p, "u_s0"), grass_s0)
u_f(gl_uniform(p, "u_d0"), grass_d0)
u_f(gl_uniform(p, "u_radius"), grass_radius)
gl_uniform1i(gl_uniform(p, "u_dbg"), grass_dbg)
var orthotex = ter_ortho_tex
var oon = F_ONE
if orthotex == 0 { orthotex = ter_height_tex; oon = F_ZERO }
r3d_bind_2d(p, "u_ortho", 4, orthotex)
u_f(gl_uniform(p, "u_ortho_on"), oon)
var lake = fl(-100000.0)
if ter_lake_ex != 0 { lake = ter_lake_level }
u_f(gl_uniform(p, "u_lake_level"), lake)
u_f(gl_uniform(p, "u_snow_line"), ter_snow_line)
sky_bind_lighting(p)
shadow_bind(p)
fog_bind(p)
u_f(gl_uniform(p, "u_spec_scale"), fl(0.15))
gl_disable(GL_CULL_FACE)
grass_draws = 0
gl_bind_vertex_array(grass_mesh.vao)
grass_tiles(16, F_ZERO, fi(300))
grass_tiles(64, fi(300), fi(1200))
grass_tiles(256, fi(1200), grass_radius)
gl_enable(GL_CULL_FACE)
}

View file

@ -0,0 +1,150 @@
# ============================================================================
# mesh.ludic — vertex data on the GPU: a Mesh record (VAO + buffers + a draw
# call), and the procedural meshes the renderer needs (a grid for the terrain,
# a full-screen triangle, a unit quad for instanced cards).
# ============================================================================
property Mesh {
vao: int = 0,
vbo: int = 0,
ebo: int = 0,
count: int = 0, # indices (ebo != 0) or vertices
mode: int = 4, # GL_TRIANGLES
itype: int = 0x1405 # GL_UNSIGNED_INT
}
function mesh_draw(m: Mesh) -> void {
gl_bind_vertex_array(m.vao)
if m.ebo != 0 { gl_draw_elements(m.mode, m.count, m.itype, null) }
else { gl_draw_arrays(m.mode, 0, m.count) }
}
function mesh_draw_instanced(m: Mesh, n: int) -> void {
gl_bind_vertex_array(m.vao)
if m.ebo != 0 { gl_draw_elements_instanced(m.mode, m.count, m.itype, null, n) }
else { gl_draw_arrays_instanced(m.mode, 0, m.count, n) }
}
# A flat n x n vertex grid over [-half, half]^2 in x/z, y = 0. Attribute 0 = (x, z).
# The terrain vertex shader lifts it with the height map.
function mesh_grid(n: int, half: int) -> Mesh {
let m = new Mesh
m.vao = gl_vao()
let nv = n * n
let v = gl_floats(nv * 2)
var k = 0
for j in 0 .. n {
for i in 0 .. n {
let x = f_sub(f_mul(f_mul(fr(i, n - 1), F_TWO), half), half)
let z = f_sub(f_mul(f_mul(fr(j, n - 1), F_TWO), half), half)
gl_put_bits(v, k, x); gl_put_bits(v, k + 1, z)
k += 2
}
}
m.vbo = gl_buffer()
gl_bind_buffer(GL_ARRAY_BUFFER, m.vbo)
gl_buffer_data(GL_ARRAY_BUFFER, gl_bytes_of(nv * 2), v, GL_STATIC_DRAW)
gl_enable_vertex_attrib_array(0)
gl_vertex_attrib_pointer(0, 2, GL_FLOAT, 0, 8, null)
free(v)
let ni = (n - 1) * (n - 1) * 6
let idx = words(ni)
k = 0
for j in 0 .. n - 1 {
for i in 0 .. n - 1 {
let a = j * n + i
idx[k] = a; idx[k + 1] = a + n; idx[k + 2] = a + 1
idx[k + 3] = a + 1; idx[k + 4] = a + n; idx[k + 5] = a + n + 1
k += 6
}
}
m.ebo = gl_buffer()
gl_bind_buffer(GL_ELEMENT_ARRAY_BUFFER, m.ebo)
gl_buffer_data(GL_ELEMENT_ARRAY_BUFFER, ni * 4, idx, GL_STATIC_DRAW)
free(idx)
m.count = ni
gl_bind_vertex_array(0)
return m
}
# The grid as patches (4 control points per cell) for tessellation shaders.
function mesh_grid_patches(n: int, half: int) -> Mesh {
let m = mesh_grid(n, half)
# rebuild the index buffer as quads
let nq = (n - 1) * (n - 1) * 4
let idx = words(nq)
var k = 0
for j in 0 .. n - 1 {
for i in 0 .. n - 1 {
let a = j * n + i
idx[k] = a; idx[k + 1] = a + 1; idx[k + 2] = a + n + 1; idx[k + 3] = a + n
k += 4
}
}
gl_bind_vertex_array(m.vao)
gl_bind_buffer(GL_ELEMENT_ARRAY_BUFFER, m.ebo)
gl_buffer_data(GL_ELEMENT_ARRAY_BUFFER, nq * 4, idx, GL_STATIC_DRAW)
gl_bind_vertex_array(0)
free(idx)
m.count = nq
m.mode = GL_PATCHES
return m
}
# A full-screen triangle with no attributes (the vertex shader uses gl_VertexID).
function mesh_fullscreen() -> Mesh {
let m = new Mesh
m.vao = gl_vao()
gl_bind_vertex_array(0)
m.count = 3
return m
}
# A unit quad in x/y ([-0.5, 0.5] x [0, 1]) with uv, attribute 0 = xy, 1 = uv.
function mesh_card() -> Mesh {
let m = new Mesh
m.vao = gl_vao()
let v = gl_floats(16)
gl_put(v, 0, -0.5); gl_put(v, 1, 0.0); gl_put(v, 2, 0.0); gl_put(v, 3, 0.0)
gl_put(v, 4, 0.5); gl_put(v, 5, 0.0); gl_put(v, 6, 1.0); gl_put(v, 7, 0.0)
gl_put(v, 8, 0.5); gl_put(v, 9, 1.0); gl_put(v, 10, 1.0); gl_put(v, 11, 1.0)
gl_put(v, 12, -0.5); gl_put(v, 13, 1.0); gl_put(v, 14, 0.0); gl_put(v, 15, 1.0)
m.vbo = gl_buffer()
gl_bind_buffer(GL_ARRAY_BUFFER, m.vbo)
gl_buffer_data(GL_ARRAY_BUFFER, 64, v, GL_STATIC_DRAW)
gl_enable_vertex_attrib_array(0)
gl_vertex_attrib_pointer(0, 2, GL_FLOAT, 0, 16, null)
gl_enable_vertex_attrib_array(1)
gl_vertex_attrib_pointer(1, 2, GL_FLOAT, 0, 16, gl_ptr(null, 8))
free(v)
let idx = words(6)
idx[0] = 0; idx[1] = 1; idx[2] = 2; idx[3] = 0; idx[4] = 2; idx[5] = 3
m.ebo = gl_buffer()
gl_bind_buffer(GL_ELEMENT_ARRAY_BUFFER, m.ebo)
gl_buffer_data(GL_ELEMENT_ARRAY_BUFFER, 24, idx, GL_STATIC_DRAW)
free(idx)
m.count = 6
gl_bind_vertex_array(0)
return m
}
# Attach a per-instance float buffer (n floats per instance, split into vec4
# attributes from `first_attr`) to a mesh's VAO. Returns the buffer id.
function mesh_instance_buffer(m: Mesh, first_attr: int, floats_per: int, data: pointer, count: int) -> int {
gl_bind_vertex_array(m.vao)
let b = gl_buffer()
gl_bind_buffer(GL_ARRAY_BUFFER, b)
gl_buffer_data(GL_ARRAY_BUFFER, gl_bytes_of(floats_per * count), data, GL_STATIC_DRAW)
var a = 0
var off = 0
while off < floats_per {
var sz = floats_per - off
if sz > 4 { sz = 4 }
gl_enable_vertex_attrib_array(first_attr + a)
gl_vertex_attrib_pointer(first_attr + a, sz, GL_FLOAT, 0, floats_per * 4, gl_ptr(null, off * 4))
gl_vertex_attrib_divisor(first_attr + a, 1)
a += 1
off += 4
}
gl_bind_vertex_array(0)
return b
}

View file

@ -0,0 +1,335 @@
# ============================================================================
# overlay.ludic — 2D drawing over the finished frame, in screen pixels with the
# origin top-left (the same space Input.mouse_x/y report): filled rectangles,
# textured quads and text from a baked font atlas (tools/blender/font_build.py).
# Quads are batched into one buffer for the whole frame and uploaded ONCE at ov_end;
# a texture change only closes a draw range. (On Apple's GL every glBufferSubData
# flushes the context and waits for the GPU; uploading per texture change made a HUD
# with dozens of changes wait dozens of times a frame: 26 ms -> 40 ms, sampled.)
# ============================================================================
const OV_MAX_QUADS: int = 6000
const OV_FLOATS: int = 8 # x, y, u, v, r, g, b, a
var ov_prog: int = 0
var ov_vao: int = 0
var ov_vbo: int = 0
var ov_buf: pointer = null
var ov_n: int = 0
var ov_tex: int = 0
var ov_white: int = 0
var ov_font: int = 0
var ov_font_adv: words = null # 95 float bits, em units
var ov_font_cols: int = 16
var ov_font_rows: int = 6
var ov_font_cell: int = 128 # px per cell in the atlas
var ov_font_em: int = 100 # px per em in the atlas
var ov_pad_x: int = 0 # float bits, em
var ov_base_y: int = 0
var ov_ready: bool = false
var ov_open: bool = false
var ov_dbg: bool = false
const OV_MAX_RANGES: int = 512
const OV_RANGE_W: int = 7 # texture, first quad, quad count, clip x, y, w, h (w = 0: none)
var ov_ranges: words = null
var ov_nr: int = 0
var ov_range_start: int = 0
var ov_clip_x: int = 0 # the current clip rectangle in screen pixels (top-left origin)
var ov_clip_y: int = 0
var ov_clip_w: int = 0
var ov_clip_h: int = 0
function overlay_init(font_dir: string) -> bool {
ov_prog = gl_program("#version 410 core\n" + r3d_shader_file("overlay.vert"), "#version 410 core\n" + r3d_shader_file("overlay.frag"))
if ov_prog == 0 { print("overlay: program failed"); return false }
ov_vao = gl_vao()
gl_bind_vertex_array(ov_vao)
ov_vbo = gl_buffer()
gl_bind_buffer(GL_ARRAY_BUFFER, ov_vbo)
gl_buffer_data(GL_ARRAY_BUFFER, gl_bytes_of(OV_MAX_QUADS * 6 * OV_FLOATS), null, GL_DYNAMIC_DRAW)
gl_enable_vertex_attrib_array(0); gl_vertex_attrib_pointer(0, 2, GL_FLOAT, 0, OV_FLOATS * 4, null)
gl_enable_vertex_attrib_array(1); gl_vertex_attrib_pointer(1, 2, GL_FLOAT, 0, OV_FLOATS * 4, gl_ptr(null, 8))
gl_enable_vertex_attrib_array(2); gl_vertex_attrib_pointer(2, 4, GL_FLOAT, 0, OV_FLOATS * 4, gl_ptr(null, 16))
gl_bind_vertex_array(0)
ov_buf = gl_floats(OV_MAX_QUADS * 6 * OV_FLOATS)
ov_ranges = words(OV_MAX_RANGES * OV_RANGE_W)
ov_white = tex_solid(255, 255, 255, 255)
# the font
ov_font_adv = words(95)
for i in 0 .. 95 { ov_font_adv[i] = fl(0.6) }
ov_pad_x = fl(0.14); ov_base_y = fl(0.30)
let meta = Fs.read_text(font_dir + "/font.json")
if meta != null {
let j = Json.parse(meta)
ov_font_cols = jint(j, "cols", 16); ov_font_rows = jint(j, "rows", 6)
ov_font_cell = f_to_int(jnum(value_get(j, "cell"))); ov_font_em = f_to_int(jnum(value_get(j, "em")))
ov_pad_x = jnum(value_get(j, "pad_x")); ov_base_y = jnum(value_get(j, "base_y"))
let adv = value_get(j, "adv")
for i in 0 .. 95 { if i < value_count(adv) { ov_font_adv[i] = jnum(value_at(adv, i)) } }
ov_font = tex_load_ex(font_dir + "/font.png", false, 0)
}
if ov_font == 0 { print("overlay: no font atlas, text disabled") }
ov_dbg = Os.has_env("R3D_FONTDBG")
ov_ready = true
return true
}
# start drawing onto the screen: blending on, depth off
function ov_begin() -> void {
if not ov_ready { return }
gl_bind_framebuffer(GL_FRAMEBUFFER, gl_screen_fbo())
gl_viewport(0, 0, gl_w, gl_h)
gl_disable(GL_DEPTH_TEST)
gl_disable(GL_CULL_FACE)
gl_enable(GL_BLEND)
gl_blend_func(GL_SRC_ALPHA, GL_ONE_MINUS_SRC_ALPHA)
gl_use_program(ov_prog)
u_f2(gl_uniform(ov_prog, "u_screen"), fi(gl_w), fi(gl_h))
ov_n = 0; ov_nr = 0; ov_range_start = 0
ov_tex = ov_white
ov_clip_w = 0
ov_open = true
}
# everything drawn until ov_unclip stays inside this rectangle (a scrolling list)
function ov_clip(x: int, y: int, w: int, h: int) -> void {
ov_close_range()
ov_clip_x = x; ov_clip_y = y; ov_clip_w = w; ov_clip_h = h
}
function ov_unclip() -> void { ov_close_range(); ov_clip_w = 0 }
# close the current draw range (quads since its start, with the current texture)
function ov_close_range() -> void {
let n = ov_n - ov_range_start
if n <= 0 { return }
if ov_nr >= OV_MAX_RANGES { ov_flush(); return }
let o = ov_nr * OV_RANGE_W
ov_ranges[o] = ov_tex; ov_ranges[o + 1] = ov_range_start; ov_ranges[o + 2] = n
ov_ranges[o + 3] = ov_clip_x; ov_ranges[o + 4] = ov_clip_y; ov_ranges[o + 5] = ov_clip_w; ov_ranges[o + 6] = ov_clip_h
ov_nr += 1
ov_range_start = ov_n
}
# one upload of everything batched so far, then a draw per range
function ov_flush() -> void {
ov_close_range()
if ov_n == 0 { ov_nr = 0; ov_range_start = 0; return }
gl_use_program(ov_prog)
gl_bind_vertex_array(ov_vao)
gl_bind_buffer(GL_ARRAY_BUFFER, ov_vbo)
gl_buffer_data(GL_ARRAY_BUFFER, gl_bytes_of(ov_n * 6 * OV_FLOATS), ov_buf, GL_STREAM_DRAW)
var last = -1
var clipped = false
for i in 0 .. ov_nr {
let o = i * OV_RANGE_W
let t = ov_ranges[o]
if t != last {
r3d_bind_2d(ov_prog, "u_tex", 0, t)
var is_font = F_ZERO
if t == ov_font { is_font = F_ONE }
u_f(gl_uniform(ov_prog, "u_is_font"), is_font)
last = t
}
if ov_ranges[o + 5] > 0 {
if not clipped { gl_enable(GL_SCISSOR_TEST); clipped = true }
gl_scissor(ov_ranges[o + 3], gl_h - ov_ranges[o + 4] - ov_ranges[o + 6], ov_ranges[o + 5], ov_ranges[o + 6])
} else if clipped { gl_disable(GL_SCISSOR_TEST); clipped = false }
gl_draw_arrays(GL_TRIANGLES, ov_ranges[o + 1] * 6, ov_ranges[o + 2] * 6)
}
if clipped { gl_disable(GL_SCISSOR_TEST) }
gl_bind_vertex_array(0)
ov_n = 0; ov_nr = 0; ov_range_start = 0
}
function ov_end() -> void {
if not ov_open { return }
ov_flush()
gl_disable(GL_BLEND)
gl_enable(GL_DEPTH_TEST)
ov_open = false
}
function ov_use_tex(t: int) -> void {
if t != ov_tex { ov_close_range(); ov_tex = t }
}
# one vertex into the batch
function ov_vert(k: int, x: int, y: int, u: int, v: int, r: int, g: int, b: int, a: int) -> void {
let o = k * OV_FLOATS
gl_put_bits(ov_buf, o, x); gl_put_bits(ov_buf, o + 1, y)
gl_put_bits(ov_buf, o + 2, u); gl_put_bits(ov_buf, o + 3, v)
gl_put_bits(ov_buf, o + 4, r); gl_put_bits(ov_buf, o + 5, g); gl_put_bits(ov_buf, o + 6, b); gl_put_bits(ov_buf, o + 7, a)
}
# a textured quad, float-bit pixel corners and uvs
function ov_quad(x0: int, y0: int, x1: int, y1: int, u0: int, v0: int, u1: int, v1: int, r: int, g: int, b: int, a: int) -> void {
if ov_n >= OV_MAX_QUADS { ov_flush() }
let k = ov_n * 6
ov_vert(k, x0, y0, u0, v0, r, g, b, a)
ov_vert(k + 1, x1, y0, u1, v0, r, g, b, a)
ov_vert(k + 2, x1, y1, u1, v1, r, g, b, a)
ov_vert(k + 3, x0, y0, u0, v0, r, g, b, a)
ov_vert(k + 4, x1, y1, u1, v1, r, g, b, a)
ov_vert(k + 5, x0, y1, u0, v1, r, g, b, a)
ov_n += 1
}
# a filled rectangle at integer pixels; colour as float bits 0..1
function ov_rect(x: int, y: int, w: int, h: int, r: int, g: int, b: int, a: int) -> void {
ov_use_tex(ov_white)
ov_quad(fi(x), fi(y), fi(x + w), fi(y + h), F_ZERO, F_ZERO, F_ONE, F_ONE, r, g, b, a)
}
function ov_frame(x: int, y: int, w: int, h: int, t: int, r: int, g: int, b: int, a: int) -> void {
ov_rect(x, y, w, t, r, g, b, a)
ov_rect(x, y + h - t, w, t, r, g, b, a)
ov_rect(x, y, t, h, r, g, b, a)
ov_rect(x + w - t, y, t, h, r, g, b, a)
}
# a whole texture at integer pixels
function ov_image(tex: int, x: int, y: int, w: int, h: int, a: int) -> void {
ov_use_tex(tex)
ov_quad(fi(x), fi(y), fi(x + w), fi(y + h), F_ZERO, F_ZERO, F_ONE, F_ONE, F_ONE, F_ONE, F_ONE, a)
}
# the width in pixels of `s` at `size` pixels per em
function ov_text_w(size: int, s: string) -> int {
var w = F_ZERO
let sp: pointer = s # bytes, not one-character strings
let n = len(sp)
for i in 0 .. n {
var c = sp[i] - 32
if c < 0 or c > 94 { c = 0 }
w = f_add(w, f_mul(ov_font_adv[c], fi(size)))
}
return f_to_int(w)
}
# text with its top-left at (x, y); returns the pen x after it
function ov_text(x: int, y: int, size: int, s: string, r: int, g: int, b: int, a: int) -> int {
if ov_font == 0 { return x }
ov_use_tex(ov_font)
let k = fr(size, ov_font_em) # atlas px -> screen px
let cell = f_mul(fi(ov_font_cell), k)
var pen = fi(x)
let base = f_add(fi(y), f_mul(fi(size), fl(0.80)))
let px = f_mul(f_mul(ov_pad_x, fi(ov_font_em)), k)
let py = f_mul(f_mul(ov_base_y, fi(ov_font_em)), k)
let sp: pointer = s
let n = len(sp)
for i in 0 .. n {
var c = sp[i] - 32
if c < 0 or c > 94 { c = 0 }
if c != 0 {
let cx = c - (c / ov_font_cols) * ov_font_cols
let cy = c / ov_font_cols
let u0 = fr(cx, ov_font_cols); let u1 = fr(cx + 1, ov_font_cols)
let v0 = fr(cy, ov_font_rows); let v1 = fr(cy + 1, ov_font_rows)
let x0 = f_sub(pen, px); let y1 = f_add(base, py)
ov_quad(x0, f_sub(y1, cell), f_add(x0, cell), y1, u0, v0, u1, v1, r, g, b, a)
}
pen = f_add(pen, f_mul(ov_font_adv[c], fi(size)))
}
return f_to_int(pen)
}
# text with a soft dark shadow under it (HUD over a bright meadow)
function ov_text_sh(x: int, y: int, size: int, s: string, r: int, g: int, b: int, a: int) -> int {
let d = size / 18 + 1
ov_text(x + d, y + d, size, s, F_ZERO, F_ZERO, F_ZERO, f_mul(a, fl(0.7)))
return ov_text(x, y, size, s, r, g, b, a)
}
function ov_text_center(cx: int, y: int, size: int, s: string, r: int, g: int, b: int, a: int) -> void {
ov_text(cx - ov_text_w(size, s) / 2, y, size, s, r, g, b, a)
}
# text wrapped at `maxw` pixels on spaces; returns the y after the last line
function ov_text_wrap(x: int, y: int, size: int, maxw: int, s: string, r: int, g: int, b: int, a: int) -> int {
let sp: pointer = s
let n = len(sp)
var first = 0
var ly = y
while first < n {
var last_space = -1
var i = first
var stop = n
while i < n {
if sp[i] == 10 { stop = i; break }
if sp[i] == 32 { last_space = i }
let piece: string = s[first .. i + 1]
if ov_text_w(size, piece) > maxw and last_space > first { stop = last_space; break }
i += 1
}
let line: string = s[first .. stop]
ov_text(x, ly, size, line, r, g, b, a)
ly += size * 13 / 10
first = stop
while first < n and (sp[first] == 32 or sp[first] == 10) { first += 1 }
}
return ly
}
# ---- more shapes for a game's interface -------------------------------------------------
# a sub-rectangle of a texture (uv corners as float bits) tinted, at integer pixels
function ov_sub(tex: int, x: int, y: int, w: int, h: int, u0: int, v0: int, u1: int, v1: int, r: int, g: int, b: int, a: int) -> void {
ov_use_tex(tex)
ov_quad(fi(x), fi(y), fi(x + w), fi(y + h), u0, v0, u1, v1, r, g, b, a)
}
# a whole texture stretched by nine slices: corners `src` texels wide in a `tw` px square
# texture, drawn `dst` pixels wide, so rounded corners keep their shape at any size
function ov_nine(tex: int, tw: int, src: int, x: int, y: int, w: int, h: int, dst: int, r: int, g: int, b: int, a: int) -> void {
ov_use_tex(tex)
let s = fr(src, tw)
let xs = words(4); let ys = words(4); let us = words(4); let vs = words(4)
xs[0] = fi(x); xs[1] = fi(x + dst); xs[2] = fi(x + w - dst); xs[3] = fi(x + w)
ys[0] = fi(y); ys[1] = fi(y + dst); ys[2] = fi(y + h - dst); ys[3] = fi(y + h)
us[0] = F_ZERO; us[1] = s; us[2] = f_sub(F_ONE, s); us[3] = F_ONE
vs[0] = F_ZERO; vs[1] = s; vs[2] = f_sub(F_ONE, s); vs[3] = F_ONE
for j in 0 .. 3 {
for i in 0 .. 3 { ov_quad(xs[i], ys[j], xs[i + 1], ys[j + 1], us[i], vs[j], us[i + 1], vs[j + 1], r, g, b, a) }
}
free(xs); free(ys); free(us); free(vs)
}
# an arbitrary quad (float-bit pixel corners, clockwise from top-left) of a texture
function ov_quad4(x0: int, y0: int, x1: int, y1: int, x2: int, y2: int, x3: int, y3: int, u0: int, v0: int, u1: int, v1: int, r: int, g: int, b: int, a: int) -> void {
if ov_n >= OV_MAX_QUADS { ov_flush() }
let k = ov_n * 6
ov_vert(k, x0, y0, u0, v0, r, g, b, a)
ov_vert(k + 1, x1, y1, u1, v0, r, g, b, a)
ov_vert(k + 2, x2, y2, u1, v1, r, g, b, a)
ov_vert(k + 3, x0, y0, u0, v0, r, g, b, a)
ov_vert(k + 4, x2, y2, u1, v1, r, g, b, a)
ov_vert(k + 5, x3, y3, u0, v1, r, g, b, a)
ov_n += 1
}
# a line of thickness `t` pixels between two points (float-bit pixels)
function ov_line(x0: int, y0: int, x1: int, y1: int, t: int, r: int, g: int, b: int, a: int) -> void {
ov_use_tex(ov_white)
let dx = f_sub(x1, x0); let dy = f_sub(y1, y0)
let l = f_max(f_sqrt(f_add(f_mul(dx, dx), f_mul(dy, dy))), fl(0.001))
let nx = f_mul(f_div(f_neg(dy), l), f_mul(t, F_HALF)); let ny = f_mul(f_div(dx, l), f_mul(t, F_HALF))
ov_quad4(f_add(x0, nx), f_add(y0, ny), f_add(x1, nx), f_add(y1, ny), f_sub(x1, nx), f_sub(y1, ny), f_sub(x0, nx), f_sub(y0, ny), F_ZERO, F_ZERO, F_ONE, F_ONE, r, g, b, a)
}
# a sub-rectangle of a texture rotated by `ang` radians about its centre (cx, cy), `w` x `h` pixels
function ov_sub_rot(tex: int, cx: int, cy: int, w: int, h: int, ang: int, u0: int, v0: int, u1: int, v1: int, r: int, g: int, b: int, a: int) -> void {
ov_use_tex(tex)
let c = f_cos(ang); let s = f_sin(ang)
let hw = f_mul(fi(w), F_HALF); let hh = f_mul(fi(h), F_HALF)
let fx = fi(cx); let fy = fi(cy)
# corners: (-hw,-hh) (hw,-hh) (hw,hh) (-hw,hh) rotated
let x0 = f_add(fx, f_sub(f_mul(f_neg(hw), c), f_mul(f_neg(hh), s))); let y0 = f_add(fy, f_add(f_mul(f_neg(hw), s), f_mul(f_neg(hh), c)))
let x1 = f_add(fx, f_sub(f_mul(hw, c), f_mul(f_neg(hh), s))); let y1 = f_add(fy, f_add(f_mul(hw, s), f_mul(f_neg(hh), c)))
let x2 = f_add(fx, f_sub(f_mul(hw, c), f_mul(hh, s))); let y2 = f_add(fy, f_add(f_mul(hw, s), f_mul(hh, c)))
let x3 = f_add(fx, f_sub(f_mul(f_neg(hw), c), f_mul(hh, s))); let y3 = f_add(fy, f_add(f_mul(f_neg(hw), s), f_mul(hh, c)))
ov_quad4(x0, y0, x1, y1, x2, y2, x3, y3, u0, v0, u1, v1, r, g, b, a)
}
# a filled circle approximated by `n` wedges (float-bit centre and radius)
function ov_disc(cx: int, cy: int, rad: int, n: int, r: int, g: int, b: int, a: int) -> void {
ov_use_tex(ov_white)
let step = f_div(f_mul(F_TWO, F_PI), fi(n))
for i in 0 .. n {
let a0 = f_mul(fi(i), step); let a1 = f_add(a0, step)
let ax = f_add(cx, f_mul(f_cos(a0), rad)); let ay = f_add(cy, f_mul(f_sin(a0), rad))
let bx = f_add(cx, f_mul(f_cos(a1), rad)); let by = f_add(cy, f_mul(f_sin(a1), rad))
ov_quad4(cx, cy, ax, ay, bx, by, cx, cy, F_ZERO, F_ZERO, F_ONE, F_ONE, r, g, b, a)
}
}
# a ring: `n` segments of thickness `t`, from angle a0 for `span` radians (float bits)
function ov_arc(cx: int, cy: int, rad: int, t: int, a0: int, span: int, n: int, r: int, g: int, b: int, a: int) -> void {
let step = f_div(span, fi(n))
for i in 0 .. n {
let b0 = f_add(a0, f_mul(fi(i), step)); let b1 = f_add(b0, step)
ov_line(f_add(cx, f_mul(f_cos(b0), rad)), f_add(cy, f_mul(f_sin(b0), rad)), f_add(cx, f_mul(f_cos(b1), rad)), f_add(cy, f_mul(f_sin(b1), rad)), t, r, g, b, a)
}
}

View file

@ -0,0 +1,7 @@
# ludic.render3d — a physically based 3D renderer over Gl.* (OpenGL 4.1 core):
# HDRI sky + image-based lighting, a GPU-generated terrain with scanned PBR
# materials, cascaded shadow maps, an HDR pipeline with bloom and ACES.
package "ludic.render3d"
version "0.1.0"
kind source
provides "R3d"

View file

@ -0,0 +1,307 @@
# ============================================================================
# post.ludic — the HDR frame and what happens to it: a 16-bit float scene
# target, a mip-chain bloom (13-tap down, tent up), and the tonemap composite
# (exposure, ACES, vignette, saturation, contrast, dither) to the screen.
# ============================================================================
const BLOOM_LEVELS: int = 6
var post_hdr: Target = null
var post_ms_fbo: int = 0 # 4x multisampled scene target, resolved into post_hdr
var post_ms_samples: int = 1 # temporal AA carries the edges; R3D_MSAA=n to compare
var post_bloom: []Target = null
var post_p_down: int = 0
var post_p_up: int = 0
var post_p_tone: int = 0
var post_fs: Mesh = null
var post_exposure: int = 0
var post_bloom_strength: int = 0
var post_vignette: int = 0
var post_saturation: int = 0
var post_contrast: int = 0
var post_w: int = 0
var post_h: int = 0
var post_auto: bool = true
var post_key: int = 0 # target mean luminance after exposure (float bits)
var post_lum: words = null
var post_mips: int = 0
var post_adapt: int = 0 # smoothed exposure (float bits)
var post_exposure_max: int = 0x41A00000 # 20: the ceiling auto-exposure may reach (night lowers it)
var post_ao: Target = null
var post_ao_blur: Target = null
var post_p_ao: int = 0
var post_p_ao_blur: int = 0
var post_ao_radius: int = 0
var post_ao_intensity: int = 0
var post_ao_strength: int = 0
var post_gi_strength: int = 0x3ECCCCCD # 0.4
var post_no_gi: bool = false
var post_ldr: Target = null
var post_depth_copy: Target = null
var post_prev: Target = null # last frame's scene colour, for the SSGI bounce only
var post_scene: Target = null # this frame's scene colour before the water, for refraction
var post_frame: int = 0
var post_color: int = 0 # the HDR colour the rest of post reads # the resolved depth, copied so passes can read it while drawing into the frame
var post_p_sharp: int = 0
var post_sharpen: int = 0
var post_grain: int = 0
# the screen-sized targets go away before post_init makes them at a new size
function post_free() -> void {
if post_hdr == null { return }
target_free(post_hdr); target_free(post_ao); target_free(post_ao_blur); target_free(post_ldr)
target_free(post_depth_copy); target_free(post_prev); target_free(post_scene)
for i in 0 .. len(post_bloom) { target_free(post_bloom[i]) }
post_hdr = null
}
function post_init(w: int, h: int) -> void {
post_w = w; post_h = h
post_hdr = target_new(w, h, GL_RGBA16F, GL_RGBA, GL_HALF_FLOAT, true, GL_LINEAR)
if post_ms_samples > 1 {
post_ms_fbo = gl_framebuffer()
gl_bind_framebuffer(GL_FRAMEBUFFER, post_ms_fbo)
let ids = gl_scratch()
gl_gen_renderbuffers(1, ids)
gl_bind_renderbuffer(GL_RENDERBUFFER, ids[0])
gl_renderbuffer_storage_multisample(GL_RENDERBUFFER, post_ms_samples, GL_RGBA16F, w, h)
gl_framebuffer_renderbuffer(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_RENDERBUFFER, ids[0])
gl_gen_renderbuffers(1, ids)
gl_bind_renderbuffer(GL_RENDERBUFFER, ids[0])
gl_renderbuffer_storage_multisample(GL_RENDERBUFFER, post_ms_samples, GL_DEPTH_COMPONENT32F, w, h)
gl_framebuffer_renderbuffer(GL_FRAMEBUFFER, GL_DEPTH_ATTACHMENT, GL_RENDERBUFFER, ids[0])
let st = gl_check_framebuffer_status(GL_FRAMEBUFFER)
if st != GL_FRAMEBUFFER_COMPLETE { print(`r3d: msaa framebuffer incomplete {st}`); post_ms_fbo = 0 }
gl_bind_framebuffer(GL_FRAMEBUFFER, 0)
}
post_bloom = new []Target
var bw = w / 2; var bh = h / 2
for i in 0 .. BLOOM_LEVELS {
push(post_bloom, target_new(max(bw, 1), max(bh, 1), GL_RGBA16F, GL_RGBA, GL_HALF_FLOAT, false, GL_LINEAR))
bw = bw / 2; bh = bh / 2
}
if post_p_down == 0 {
post_p_down = r3d_program("fullscreen.vert", "bloom_down.frag", "")
post_p_up = r3d_program("fullscreen.vert", "bloom_up.frag", "")
post_p_tone = r3d_program("fullscreen.vert", "tonemap.frag", "")
}
# Full resolution, not half. The occlusion is reconstructed from depth differences,
# so on a surface seen at a grazing angle its gradient is steep in screen space; at
# half resolution that aliased into wide, screen-crossing bands which the bilinear
# upsample in the tonemapper then stretched over the whole ground. They read as thin
# transparent black bars, appear only where there is depth (never on the sky), and
# are nothing to do with the shadow map or the reflection.
post_ao = target_new(w, h, GL_RGBA16F, GL_RGBA, GL_HALF_FLOAT, false, GL_LINEAR)
post_ao_blur = target_new(w, h, GL_RGBA16F, GL_RGBA, GL_HALF_FLOAT, false, GL_LINEAR)
if post_p_ao == 0 { post_p_ao = r3d_program("fullscreen.vert", "ssgi.frag", ""); post_p_ao_blur = r3d_program("fullscreen.vert", "ssao_blur.frag", "") }
post_ldr = target_new(w, h, GL_RGBA8, GL_RGBA, GL_UNSIGNED_BYTE, false, GL_LINEAR)
post_depth_copy = target_new(w, h, GL_R8, GL_RED, GL_UNSIGNED_BYTE, true, GL_NEAREST)
post_prev = target_new(w, h, GL_RGBA16F, GL_RGBA, GL_HALF_FLOAT, false, GL_LINEAR)
post_scene = target_new(w, h, GL_RGBA16F, GL_RGBA, GL_HALF_FLOAT, false, GL_LINEAR)
if post_p_sharp == 0 { post_p_sharp = r3d_program("fullscreen.vert", "sharpen.frag", "") }
post_sharpen = fl(1.2)
post_grain = fl(0.025)
post_ao_radius = fl(0.7)
post_ao_intensity = fl(1.4)
post_ao_strength = fl(0.8)
post_fs = mesh_fullscreen()
post_exposure = fl(0.36)
post_bloom_strength = fl(0.06)
post_vignette = fl(0.35)
post_saturation = fl(1.04)
post_contrast = fl(1.12)
post_key = fl(0.19)
post_lum = words(4)
var m = 1; var sz = max(w, h)
while sz > 1 { sz = sz / 2; m += 1 }
post_mips = m
post_adapt = F_ZERO
}
# Mean scene luminance from the HDR mip chain -> exposure = key / mean, eased over
# frames. The value comes back through a pixel buffer one frame late: a direct
# glGetTexImage waits for the GPU to finish the whole frame, which serialised the
# CPU and the GPU. With the fly-camera demo that cost little (the CPU had nothing
# else to do); with the game's animals, HUD and rules on the CPU it doubled the frame
# (60 ms -> 28 ms when the read went asynchronous, measured 2026-09-09).
# ... and even that asynchronous read blocked on Apple's GL (glGetTexImage into a pixel
# buffer still synchronised the texture: 50% of the CPU's frame waiting, sampled), so
# the adaptation now stays on the GPU: a 1x1 pass (adapt.frag) eases last frame's value
# toward key / mean and the tonemapper samples it. The CPU never waits for the picture.
var post_adapt_t: []Target = null
var post_adapt_i: int = 0
var post_p_adapt: int = 0
var post_adapt_reset: bool = true
function post_measure() -> void {
gl_bind_texture(GL_TEXTURE_2D, post_hdr.color)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_MIN_FILTER, GL_LINEAR_MIPMAP_LINEAR)
gl_generate_mipmap(GL_TEXTURE_2D)
if post_adapt_t == null {
post_adapt_t = new []Target
for k in 0 .. 2 { push(post_adapt_t, target_new(1, 1, GL_R32F, GL_RED, GL_FLOAT, false, GL_NEAREST)) }
post_adapt_reset = true
}
if post_p_adapt == 0 { post_p_adapt = r3d_program("fullscreen.vert", "adapt.frag", "") }
let next = 1 - post_adapt_i
target_bind(post_adapt_t[next])
gl_disable(GL_DEPTH_TEST)
gl_use_program(post_p_adapt)
r3d_bind_2d(post_p_adapt, "u_scene", 0, post_hdr.color)
r3d_bind_2d(post_p_adapt, "u_prev", 1, post_adapt_t[post_adapt_i].color)
u_f(gl_uniform(post_p_adapt, "u_lod"), fi(post_mips - 1))
u_f(gl_uniform(post_p_adapt, "u_key"), post_key)
u_f(gl_uniform(post_p_adapt, "u_max"), post_exposure_max)
u_f(gl_uniform(post_p_adapt, "u_rate"), fl(0.08))
var reset = F_ZERO
if post_adapt_reset { reset = F_ONE; post_adapt_reset = false }
u_f(gl_uniform(post_p_adapt, "u_reset"), reset)
mesh_draw(post_fs)
post_adapt_i = next
gl_bind_texture(GL_TEXTURE_2D, post_hdr.color)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_MIN_FILTER, GL_LINEAR)
}
function post_begin_scene() -> void {
target_bind(post_hdr)
if post_ms_fbo != 0 { gl_bind_framebuffer(GL_FRAMEBUFFER, post_ms_fbo); gl_enable(GL_MULTISAMPLE) }
gl_enable(GL_DEPTH_TEST)
gl_depth_func(GL_LESS)
gl_depth_mask(1)
gl_enable(GL_CULL_FACE)
gl_cull_face(GL_BACK)
gl_clear_color(0.0, 0.0, 0.0, 1.0)
gl_clear(GL_COLOR_BUFFER_BIT | GL_DEPTH_BUFFER_BIT)
}
# resolve the multisampled scene into the plain HDR target (colour + depth)
function post_resolve() -> void {
if post_ms_fbo != 0 {
gl_bind_framebuffer(GL_READ_FRAMEBUFFER, post_ms_fbo)
gl_bind_framebuffer(GL_DRAW_FRAMEBUFFER, post_hdr.fbo)
gl_blit_framebuffer(0, 0, post_w, post_h, 0, 0, post_w, post_h, GL_COLOR_BUFFER_BIT | GL_DEPTH_BUFFER_BIT, GL_NEAREST)
}
# the depth copy every pass after this may read while the frame is still being drawn into
gl_bind_framebuffer(GL_READ_FRAMEBUFFER, post_hdr.fbo)
gl_bind_framebuffer(GL_DRAW_FRAMEBUFFER, post_depth_copy.fbo)
gl_blit_framebuffer(0, 0, post_w, post_h, 0, 0, post_w, post_h, GL_DEPTH_BUFFER_BIT, GL_NEAREST)
gl_bind_framebuffer(GL_FRAMEBUFFER, 0)
}
# There is no temporal anti-aliasing. It was reprojecting every pixel through the
# scene depth, which on water is the surface plane while the pixel's content is the
# reflection behind it — so the mirror image was fetched from the wrong place and, at
# 0.92 history, dragged several frames behind the camera as it turned. Geometry edges
# and the alpha-tested vegetation are covered by the 4x MSAA + alpha-to-coverage the
# scene already renders with, and the projection is no longer jittered, so nothing is
# left needing a temporal resolve.
# The lake bed, as drawn, before any water goes over it. Water reads this to refract and
# then absorb it, which is what makes the surface read as a body of water rather than a
# sheet laid over the ground: the bottom is seen THROUGH the water, tinted and dimmed by
# how far the light travelled, instead of being the dry terrain showing through an alpha.
function post_capture_scene() -> void {
gl_bind_framebuffer(GL_READ_FRAMEBUFFER, post_hdr.fbo)
gl_bind_framebuffer(GL_DRAW_FRAMEBUFFER, post_scene.fbo)
gl_blit_framebuffer(0, 0, post_w, post_h, 0, 0, post_w, post_h, GL_COLOR_BUFFER_BIT, GL_NEAREST)
gl_bind_framebuffer(GL_FRAMEBUFFER, post_hdr.fbo)
gl_viewport(0, 0, post_w, post_h)
}
# Keep a copy of the finished scene colour: the SSGI bounce reads last frame's colour.
function post_capture_prev() -> void {
gl_bind_framebuffer(GL_READ_FRAMEBUFFER, post_hdr.fbo)
gl_bind_framebuffer(GL_DRAW_FRAMEBUFFER, post_prev.fbo)
gl_blit_framebuffer(0, 0, post_w, post_h, 0, 0, post_w, post_h, GL_COLOR_BUFFER_BIT, GL_NEAREST)
gl_bind_framebuffer(GL_FRAMEBUFFER, 0)
post_frame += 1
}
function post_ssao_pass() -> void {
gl_disable(GL_DEPTH_TEST)
gl_disable(GL_BLEND)
target_bind(post_ao)
gl_use_program(post_p_ao)
r3d_bind_2d(post_p_ao, "u_depth", 0, post_hdr.depth)
r3d_bind_2d(post_p_ao, "u_prev_color", 1, post_prev.color)
u_f(gl_uniform(post_p_ao, "u_frame"), fi(post_frame % 64))
u_mat4(gl_uniform(post_p_ao, "u_inv_proj"), cam_inv_proj)
u_mat4(gl_uniform(post_p_ao, "u_proj"), cam_proj)
u_f2(gl_uniform(post_p_ao, "u_texel"), fr(1, post_w), fr(1, post_h))
u_f(gl_uniform(post_p_ao, "u_radius"), post_ao_radius)
u_f(gl_uniform(post_p_ao, "u_intensity"), post_ao_intensity)
mesh_draw(post_fs)
target_bind(post_ao_blur)
gl_use_program(post_p_ao_blur)
r3d_bind_2d(post_p_ao_blur, "u_ao", 0, post_ao.color)
r3d_bind_2d(post_p_ao_blur, "u_depth", 1, post_hdr.depth)
u_f2(gl_uniform(post_p_ao_blur, "u_texel"), fr(1, post_ao.w), fr(1, post_ao.h))
mesh_draw(post_fs)
}
function post_bloom_pass() -> void {
gl_disable(GL_DEPTH_TEST)
gl_disable(GL_BLEND)
var src = post_color
var sw = post_w; var sh = post_h
gl_use_program(post_p_down)
for i in 0 .. BLOOM_LEVELS {
let t = post_bloom[i]
target_bind(t)
r3d_bind_2d(post_p_down, "u_src", 0, src)
u_f2(gl_uniform(post_p_down, "u_texel"), fr(1, sw), fr(1, sh))
var th = f_neg1()
if i == 0 { th = fl(1.2) }
u_f(gl_uniform(post_p_down, "u_threshold"), th)
mesh_draw(post_fs)
src = t.color; sw = t.w; sh = t.h
}
gl_use_program(post_p_up)
gl_enable(GL_BLEND)
gl_blend_func(GL_ONE, GL_ONE)
var i = BLOOM_LEVELS - 1
while i > 0 {
let from = post_bloom[i]
let to = post_bloom[i - 1]
target_bind(to)
r3d_bind_2d(post_p_up, "u_src", 0, from.color)
u_f2(gl_uniform(post_p_up, "u_texel"), fr(1, from.w), fr(1, from.h))
u_f(gl_uniform(post_p_up, "u_radius"), F_ONE)
mesh_draw(post_fs)
i -= 1
}
gl_disable(GL_BLEND)
}
function post_tonemap(color_tex: int) -> void {
if post_auto { post_measure() }
target_bind(post_ldr)
gl_disable(GL_DEPTH_TEST)
gl_use_program(post_p_tone)
r3d_bind_2d(post_p_tone, "u_hdr", 0, color_tex)
r3d_bind_2d(post_p_tone, "u_bloom", 1, post_bloom[0].color)
r3d_bind_2d(post_p_tone, "u_ao", 2, post_ao_blur.color)
u_f(gl_uniform(post_p_tone, "u_ao_strength"), post_ao_strength)
u_f(gl_uniform(post_p_tone, "u_gi_strength"), post_gi_strength)
u_f(gl_uniform(post_p_tone, "u_exposure"), post_exposure)
var auto = F_ZERO
if post_auto and post_adapt_t != null { auto = F_ONE; r3d_bind_2d(post_p_tone, "u_adapt", 3, post_adapt_t[post_adapt_i].color) }
u_f(gl_uniform(post_p_tone, "u_auto"), auto)
u_f(gl_uniform(post_p_tone, "u_bloom_strength"), post_bloom_strength)
u_f(gl_uniform(post_p_tone, "u_vignette"), post_vignette)
u_f(gl_uniform(post_p_tone, "u_saturation"), post_saturation)
u_f(gl_uniform(post_p_tone, "u_contrast"), post_contrast)
u_f3(gl_uniform(post_p_tone, "u_wb"), fl(1.02), F_ONE, fl(0.97))
u_f3(gl_uniform(post_p_tone, "u_lift"), fl(0.004), fl(0.004), fl(0.012))
u_f3(gl_uniform(post_p_tone, "u_gain"), fl(0.99), fl(0.995), fl(1.0))
mesh_draw(post_fs)
# sharpen + grain onto the screen
gl_bind_framebuffer(GL_FRAMEBUFFER, gl_screen)
gl_viewport(0, 0, gl_w, gl_h)
gl_use_program(post_p_sharp)
r3d_bind_2d(post_p_sharp, "u_src", 0, post_ldr.color)
u_f2(gl_uniform(post_p_sharp, "u_texel"), fr(1, post_w), fr(1, post_h))
u_f(gl_uniform(post_p_sharp, "u_amount"), post_sharpen)
u_f(gl_uniform(post_p_sharp, "u_grain"), post_grain)
u_f(gl_uniform(post_p_sharp, "u_time"), r3d_time)
mesh_draw(post_fs)
}

View file

@ -0,0 +1,385 @@
# ============================================================================
# prof.ludic — per-pass GPU timing (R3D_PROF=1).
#
# Wall-clock timing of this renderer is useless at pass granularity: run-to-run
# variance on a laptop GPU is ±15%, which is larger than most passes. GL_TIME_ELAPSED
# queries measure what the GPU actually spent inside each pass, and averaging a few
# hundred frames inside ONE run cancels the run-to-run noise entirely.
#
# Each slot owns a small ring of query objects. A query is read back only after
# enough frames have passed that its result is certainly available, so profiling
# never introduces the very stall it is trying to measure.
# ============================================================================
const PROF_SLOTS: int = 16
const PROF_RING: int = 4 # frames of latency before a result is read
var prof_on: bool = false
var prof_names: []pointer = null
var prof_ids: words = null # PROF_SLOTS * PROF_RING query objects
var prof_ns: []long = null # accumulated nanoseconds per slot
var prof_hits: []long = null # samples accumulated per slot
var prof_n: int = 0 # slots in use
var prof_frame: int = 0
var prof_active: int = -1 # slot whose query is currently open
var prof_scratch: words = null
function prof_init() -> void {
prof_on = Os.has_env("R3D_PROF")
if not prof_on { return }
prof_names = new []pointer
prof_ns = new []long
prof_hits = new []long
prof_ids = words(PROF_SLOTS * PROF_RING)
prof_scratch = words(4)
gl_gen_queries(PROF_SLOTS * PROF_RING, prof_ids)
prof_n = 0
prof_frame = 0
prof_active = -1
}
# the slot index for `name`, registering it on first sight (order = frame order)
function prof_slot(name: pointer) -> int {
var i = 0
while i < prof_n {
if prof_names[i] == name { return i }
i += 1
}
if prof_n >= PROF_SLOTS { return -1 }
push(prof_names, name)
push(prof_ns, 0)
push(prof_hits, 0)
prof_n += 1
return prof_n - 1
}
function prof_begin(name: pointer) -> void {
if not prof_on { return }
if prof_active >= 0 { return } # GL_TIME_ELAPSED queries cannot nest
let s = prof_slot(name)
if s < 0 { return }
prof_active = s
gl_begin_query(GL_TIME_ELAPSED, prof_ids[s * PROF_RING + (prof_frame % PROF_RING)])
}
function prof_end() -> void {
if not prof_on { return }
if prof_active < 0 { return }
gl_end_query(GL_TIME_ELAPSED)
prof_active = -1
}
# Collect the queries issued PROF_RING-1 frames ago — long finished, so no stall.
function prof_collect() -> void {
if not prof_on { return }
prof_frame += 1
if prof_frame < PROF_RING { return }
let slot_frame = (prof_frame + 1) % PROF_RING
var s = 0
while s < prof_n {
let q = prof_ids[s * PROF_RING + slot_frame]
gl_get_query_objectiv(q, GL_QUERY_RESULT_AVAILABLE, prof_scratch)
if prof_scratch[0] != 0 {
gl_get_query_objectui64v(q, GL_QUERY_RESULT, prof_scratch)
# the low 32 bits are ample: a pass is far below 4 seconds
prof_ns[s] = prof_ns[s] + prof_scratch[0]
prof_hits[s] = prof_hits[s] + 1
}
s += 1
}
}
function prof_report() -> void {
if not prof_on { return }
print("")
print("GPU time per pass (mean over the run):")
var total = 0
var s = 0
while s < prof_n {
if prof_hits[s] > 0 { total = total + prof_ns[s] / prof_hits[s] }
s += 1
}
s = 0
while s < prof_n {
if prof_hits[s] > 0 {
let us = prof_ns[s] / prof_hits[s] / 1000
var pct = 0
if total > 0 { pct = (prof_ns[s] / prof_hits[s]) * 100 / total }
print(` {Text.pad_right(prof_names[s], 22)} {Text.pad_left(string(us), 7)} us {string(pct)}%`)
}
s += 1
}
print(` {Text.pad_right("TOTAL", 22)} {Text.pad_left(string(total / 1000), 7)} us`)
}
# ---- streaming work per frame (R3D_PROF=1) ----------------------------------
# Stutter while walking is not visible in an average: the cover for a newly entered
# chunk is generated in whichever frame the camera crosses a 32 m cell, so one frame
# in fifty does all the work. `Time.delta` here is a fixed 60 Hz timestep and there is
# no finer wall clock in the runtime, so measure the WORK instead — instances generated
# per frame is exactly what the hitch is made of, and it needs no clock at all.
var prof_gen: []int = null
var prof_gen_cur: int = 0
var prof_ft: []long = null # real frame times, microseconds
var prof_last_us: long = 0
function prof_gen_add(n: int) -> void {
if not prof_on { return }
prof_gen_cur += n
}
# Per-chunk generation, so a hitch can be pinned on a stream and a band rather than on
# "streaming". One chunk is generated atomically, so the worst chunk is the worst frame.
var prof_chunk_us: []long = null
var prof_chunk_kind: []int = null
var prof_chunk_band: []int = null
var prof_chunk_n: []int = null
var prof_gen_us_cur: long = 0
var prof_gen_us: []long = null # generation microseconds per frame
# CPU work that only happens on some frames — which is what a hitch is made of.
var prof_layer_us_cur: long = 0 # rebuilding and uploading instance buffers
var prof_bake_us_cur: long = 0 # the height-field shadow rebake
var prof_layer_us: []long = null
var prof_bake_us: []long = null
var prof_dt: []long = null # frame time, aligned with the arrays above
var prof_up_cur: long = 0 # instance bytes uploaded this frame
var prof_up: []long = null
# Where a frame's time went at the coarsest useful split: work this process did, and
# time spent waiting for the GPU to finish it. Headless drains the GPU inside swap, so
# the two are cleanly separable there.
var prof_pre_swap: long = 0
var prof_cpu_us: []long = null
function prof_before_swap() -> void { if prof_on { prof_pre_swap = gl_now_us() } }
# The single most expensive CPU phase of each frame, and what it was. A hitch is one
# phase running long on one frame, so recording the worst one per frame is enough to
# name it without keeping a timeline.
var prof_mark_last: long = 0
var prof_mark_best: long = 0
var prof_mark_name: pointer = null
var prof_mark_us: []long = null
var prof_mark_who: []pointer = null
function prof_mark_start() -> void { if prof_on { prof_mark_last = gl_now_us(); prof_mark_best = 0; prof_mark_name = null } }
function prof_cpu_mark(name: pointer) -> void {
if not prof_on { return }
let now = gl_now_us()
let d = now - prof_mark_last
prof_mark_last = now
if d > prof_mark_best { prof_mark_best = d; prof_mark_name = name }
}
# the game's own per-frame work, kept apart from the renderer's
var prof_game_us_cur: long = 0
var prof_game_us: []long = null
function prof_game_add(us: long) -> void { if prof_on { prof_game_us_cur = prof_game_us_cur + us } }
function prof_layer_add(us: long, bytes: long) -> void {
if not prof_on { return }
prof_layer_us_cur = prof_layer_us_cur + us
prof_up_cur = prof_up_cur + bytes
}
function prof_bake_add(us: long) -> void { if prof_on { prof_bake_us_cur = prof_bake_us_cur + us } }
function prof_chunk(kind: int, band: int, count: int, us: long) -> void {
if not prof_on { return }
if prof_chunk_us == null {
prof_chunk_us = new []int; prof_chunk_kind = new []int
prof_chunk_band = new []int; prof_chunk_n = new []int
}
push(prof_chunk_us, us); push(prof_chunk_kind, kind)
push(prof_chunk_band, band); push(prof_chunk_n, count)
prof_gen_us_cur += us
}
function prof_chunk_report() -> void {
if not prof_on or prof_chunk_us == null { return }
print("")
print("chunk generation (one chunk is atomic, so the worst chunk is the worst frame):")
# totals per (kind, band)
var kinds = new []int
var bands = new []int
var tot = new []long
var cnt = new []int
var mx = new []long
var i = 0
while i < len(prof_chunk_us) {
var f = -1
var j = 0
while j < len(kinds) { if kinds[j] == prof_chunk_kind[i] and bands[j] == prof_chunk_band[i] { f = j }; j += 1 }
if f < 0 {
push(kinds, prof_chunk_kind[i]); push(bands, prof_chunk_band[i])
push(tot, 0); push(cnt, 0); push(mx, 0)
f = len(kinds) - 1
}
tot[f] = tot[f] + prof_chunk_us[i]
cnt[f] = cnt[f] + 1
if prof_chunk_us[i] > mx[f] { mx[f] = prof_chunk_us[i] }
i += 1
}
var k = 0
while k < len(kinds) {
print(` kind {string(kinds[k])} band {string(bands[k])}: {string(cnt[k])} chunks, mean {string(tot[k] / cnt[k])} us, worst {string(mx[k])} us, total {string(tot[k] / 1000)} ms`)
k += 1
}
# the per-frame distribution of generation time: this is the hitch itself
if prof_gen_us == null or len(prof_gen_us) < 16 { return }
let sorted = new []long
var a = 8
while a < len(prof_gen_us) { push(sorted, prof_gen_us[a]); a += 1 }
var x = 1
while x < len(sorted) {
let v = sorted[x]
var y = x - 1
while y >= 0 and sorted[y] > v { sorted[y + 1] = sorted[y]; y -= 1 }
sorted[y + 1] = v
x += 1
}
let n = len(sorted)
var busy = 0
var t: long = 0
var z = 0
while z < n { if sorted[z] > 0 { busy += 1 }; t += sorted[z]; z += 1 }
print(` generation per frame (us): p95 {string(sorted[(n * 95) / 100])} p99 {string(sorted[(n * 99) / 100])} worst {string(sorted[n - 1])} frames that generated: {string(busy)} of {string(n)} total {string(t / 1000)} ms`)
}
function prof_gen_frame() -> void {
if not prof_on { return }
if prof_gen == null {
prof_gen = new []int; prof_ft = new []long
prof_gen_us = new []long; prof_layer_us = new []long
prof_bake_us = new []long; prof_dt = new []long; prof_up = new []long
prof_cpu_us = new []long; prof_game_us = new []long
prof_mark_us = new []long; prof_mark_who = new []pointer
}
push(prof_gen, prof_gen_cur)
prof_gen_cur = 0
let now = gl_now_us()
var dtf: long = 0
if prof_last_us != 0 { push(prof_ft, now - prof_last_us); dtf = now - prof_last_us }
prof_last_us = now
# everything the frame just ended spent on work it only does sometimes
push(prof_dt, dtf)
push(prof_gen_us, prof_gen_us_cur)
push(prof_layer_us, prof_layer_us_cur)
push(prof_bake_us, prof_bake_us_cur)
push(prof_up, prof_up_cur)
var cpu: long = 0
if prof_pre_swap != 0 and dtf != 0 { cpu = prof_pre_swap - (now - dtf) }
push(prof_cpu_us, cpu)
push(prof_game_us, prof_game_us_cur)
push(prof_mark_us, prof_mark_best)
if prof_mark_name == null { push(prof_mark_who, "-") } else { push(prof_mark_who, prof_mark_name) }
prof_gen_us_cur = 0; prof_layer_us_cur = 0; prof_bake_us_cur = 0; prof_up_cur = 0; prof_game_us_cur = 0
}
# the distribution of REAL frame times: stutter lives in the tail, not the mean
function prof_ft_report() -> void {
if not prof_on { return }
if prof_ft == null or len(prof_ft) < 16 { return }
let sorted = new []long
var i = 8
while i < len(prof_ft) { push(sorted, prof_ft[i]); i += 1 }
var a = 1
while a < len(sorted) {
let v = sorted[a]
var b = a - 1
while b >= 0 and sorted[b] > v { sorted[b + 1] = sorted[b]; b -= 1 }
sorted[b + 1] = v
a += 1
}
let n = len(sorted)
let med = sorted[n / 2]
var over = 0
var j = 0
while j < n { if sorted[j] > med * 2 { over += 1 }; j += 1 }
print("")
print("REAL frame time (us):")
print(` median {string(med)} p95 {string(sorted[(n * 95) / 100])} p99 {string(sorted[(n * 99) / 100])} worst {string(sorted[n - 1])}`)
# A hitch is not the mean moving: it is the count of frames that took noticeably
# longer than the frame before them. 1.3x median is about where it stops being smooth.
var o13 = 0
var o15 = 0
var q = 0
while q < n {
if sorted[q] * 10 > med * 13 { o13 += 1 }
if sorted[q] * 2 > med * 3 { o15 += 1 }
q += 1
}
# How much time the run spent being slower than itself: the sum of every frame's
# excess over 1.2x the median. One number that goes down when hitching goes down, and
# that a handful of unlucky frames cannot dominate the way a maximum can.
var excess: long = 0
q = 0
while q < n {
let lim = (med * 12) / 10
if sorted[q] > lim { excess = excess + (sorted[q] - lim) }
q += 1
}
print(` fps at median {string(1000000 / med)} frames over 1.3x median: {string(o13)}, over 1.5x: {string(o15)}, over 2x: {string(over)} — of {string(n)}`)
print(` stutter: {string(excess / 1000)} ms of frame time beyond 1.2x median over the run`)
print(` streaming totals over the run: walk {string(stream_us_walk / 1000)} ms (generate {string(stream_us_gen / 1000)} ms, gather {string(stream_us_gather / 1000)} ms), {string(stream_walks)} stream-walks`)
}
# The slowest frames of the run, with the once-in-a-while CPU work that landed in them.
# An average never shows a hitch; this is the list of the frames you actually felt.
function prof_hitch_report() -> void {
if not prof_on or prof_dt == null or len(prof_dt) < 32 { return }
let n = len(prof_dt)
# median, for a sense of what "slow" means here
let sorted = new []long
var i = 8
while i < n { push(sorted, prof_dt[i]); i += 1 }
var a = 1
while a < len(sorted) {
let v = sorted[a]
var b = a - 1
while b >= 0 and sorted[b] > v { sorted[b + 1] = sorted[b]; b -= 1 }
sorted[b + 1] = v
a += 1
}
let med = sorted[len(sorted) / 2]
print("")
print(`the 20 slowest frames (median {string(med / 1000)}.{string((med / 100) % 10)} ms), and what was in them:`)
print(" frame dt cpu gpu-wait cover-gen game uploaded slowest CPU phase")
var shown = 0
var cut: long = sorted[len(sorted) - 1]
while shown < 20 and cut > med {
# the next slowest frame at or below `cut`
var best = -1
var bestv: long = -1
var k = 8
while k < n {
if prof_dt[k] <= cut and prof_dt[k] > bestv { bestv = prof_dt[k]; best = k }
k += 1
}
if best < 0 { return }
print(` {Text.pad_left(string(best), 6)} {Text.pad_left(string(prof_dt[best]), 6)}us {Text.pad_left(string(prof_cpu_us[best]), 7)}us {Text.pad_left(string(prof_dt[best] - prof_cpu_us[best]), 8)}us {Text.pad_left(string(prof_gen_us[best]), 7)}us {Text.pad_left(string(prof_game_us[best]), 7)}us {Text.pad_left(string(prof_up[best] / 1024), 7)}KB {Text.pad_right(prof_mark_who[best], 18)} {Text.pad_left(string(prof_mark_us[best]), 7)}us`)
cut = bestv - 1
shown += 1
}
}
function prof_gen_report() -> void {
if not prof_on { return }
if prof_gen == null or len(prof_gen) < 8 { return }
let sorted = new []int
var i = 8 # skip the first frames: one-off initial fill
while i < len(prof_gen) { push(sorted, prof_gen[i]); i += 1 }
var a = 1
while a < len(sorted) {
let v = sorted[a]
var b = a - 1
while b >= 0 and sorted[b] > v { sorted[b + 1] = sorted[b]; b -= 1 }
sorted[b + 1] = v
a += 1
}
let n = len(sorted)
var total = 0
var busy = 0
var j = 0
while j < n { total += sorted[j]; if sorted[j] > 0 { busy += 1 }; j += 1 }
print("")
print("ground cover generated per frame (the source of walking stutter):")
print(` frames that generated anything: {string(busy)} of {string(n)}`)
print(` median {string(sorted[n / 2])} p95 {string(sorted[(n * 95) / 100])} worst {string(sorted[n - 1])} total {string(total)}`)
}

View file

@ -0,0 +1,116 @@
# ============================================================================
# programs.ludic — GLSL from files. A shader file carries no #version line; the
# loader prepends "#version 410 core", any per-variant defines, and the shared
# noise.glsl + lighting.glsl chunks for fragment stages, then compiles it.
# ============================================================================
# Where the renderer's own files are. The shaders belong to this package and ship with
# it — in the Ludic checkout they are under packages/, and in an installed toolchain
# under $LUDIC_HOME/packages/ — so a game built outside the Ludic tree does not have to
# copy them in. The scanned CC0 materials are the other half: too large to ship with a
# toolchain and redistributable from their origin, so they are fetched into the project
# (`ludic assets`) and read from there.
var r3d_root: string = "packages/ludic.render3d" # where shaders/ lives
var r3d_assets: string = "assets/polyhaven" # where the CC0 assets live
var r3d_root_found: bool = false
# defines prepended to EVERY program (set before any is built): renderer-wide switches
var r3d_global_defs: string = ""
var r3d_noise_src: string = null
var r3d_lighting_src: string = null
function r3d_set_paths(root: string, assets: string) -> void { r3d_root = root; r3d_assets = assets }
# The install root, as the compiler computes it: $LUDIC_HOME, else the directory of the
# `ludic` on PATH. A game running from its own tree finds this package there.
function r3d_home() -> string {
let env = Os.env("LUDIC_HOME")
if env != null and env != "" { return env }
return `{Os.env("HOME")}/.ludic`
}
# Settle r3d_root on first use: the package as checked out beside the project, else the
# copy that ships with the toolchain.
function r3d_find_root() -> void {
if r3d_root_found { return }
r3d_root_found = true
if Fs.exists(`{r3d_root}/shaders/lighting.glsl`) { return }
let home = r3d_home()
let alt = `{home}/packages/ludic.render3d`
if Fs.exists(`{alt}/shaders/lighting.glsl`) { r3d_root = alt; return }
print(`r3d: cannot find the renderer's shaders (looked in {r3d_root}/shaders and {alt}/shaders)`)
}
function r3d_shader_file(name: string) -> string {
r3d_find_root()
let path = `{r3d_root}/shaders/{name}`
let s = Fs.read_text(path)
if s == null { print(`r3d: missing shader {path}`); return "" }
return s
}
function r3d_shader_src(name: string, defines: string, is_frag: bool) -> string {
if r3d_noise_src == null { r3d_noise_src = r3d_shader_file("noise.glsl") }
if r3d_lighting_src == null { r3d_lighting_src = r3d_shader_file("lighting.glsl") }
var s = "#version 410 core\n" + r3d_global_defs + defines
if is_frag { s = s + r3d_noise_src + r3d_lighting_src }
return s + r3d_shader_file(name)
}
# Build a program with tessellation control/evaluation between vertex and fragment.
function r3d_program_tess(vs: string, tcs: string, tes: string, fs: string, defines: string) -> int {
let p = gl_program5(r3d_shader_src(vs, defines, false), r3d_shader_src(tcs, defines, false),
r3d_shader_src(tes, defines, false), null, r3d_shader_src(fs, defines, true))
if p == 0 { print(`r3d: tess program failed: {vs} + {tcs} + {tes} + {fs}`) }
return p
}
# Build a program from a vertex + fragment file pair (defines apply to both).
function r3d_program(vs: string, fs: string, defines: string) -> int {
let p = gl_program(r3d_shader_src(vs, defines, false), r3d_shader_src(fs, defines, true))
if p == 0 { print(`r3d: program failed: {vs} + {fs}`) }
return p
}
# Bind a texture to a unit and point a sampler uniform at it.
function r3d_bind_tex(prog: int, name: string, unit: int, target: int, tex: int) -> void {
gl_active_texture(GL_TEXTURE0 + unit)
gl_bind_texture(target, tex)
gl_uniform1i(gl_get_uniform_location(prog, name), unit)
}
function r3d_bind_2d(prog: int, name: string, unit: int, tex: int) -> void { r3d_bind_tex(prog, name, unit, GL_TEXTURE_2D, tex) }
# A framebuffer with one colour texture (and optionally a depth texture).
property Target {
fbo: int = 0,
color: int = 0,
depth: int = 0,
w: int = 0,
h: int = 0
}
function target_new(w: int, h: int, ifmt: int, fmt: int, ty: int, with_depth: bool, filter: int) -> Target {
let t = new Target
t.w = w; t.h = h
t.fbo = gl_framebuffer()
gl_bind_framebuffer(GL_FRAMEBUFFER, t.fbo)
t.color = tex_target(w, h, ifmt, fmt, ty, filter)
gl_framebuffer_texture2d(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, t.color, 0)
if with_depth {
t.depth = tex_target(w, h, GL_DEPTH_COMPONENT32F, GL_DEPTH_COMPONENT, GL_FLOAT, GL_NEAREST)
gl_framebuffer_texture2d(GL_FRAMEBUFFER, GL_DEPTH_ATTACHMENT, GL_TEXTURE_2D, t.depth, 0)
}
let st = gl_check_framebuffer_status(GL_FRAMEBUFFER)
if st != GL_FRAMEBUFFER_COMPLETE { print(`r3d: framebuffer incomplete {st} ({w}x{h})`) }
gl_bind_framebuffer(GL_FRAMEBUFFER, 0)
return t
}
function target_free(t: Target) -> void {
if t == null { return }
let ids = gl_scratch()
ids[0] = t.fbo; gl_delete_framebuffers(1, ids)
ids[0] = t.color; gl_delete_textures(1, ids)
if t.depth != 0 { ids[0] = t.depth; gl_delete_textures(1, ids) }
}
function target_bind(t: Target) -> void {
gl_bind_framebuffer(GL_FRAMEBUFFER, t.fbo)
gl_viewport(0, 0, t.w, t.h)
}

View file

@ -0,0 +1,95 @@
# ============================================================================
# quat.ludic — unit quaternions (x, y, z, w) as 4-word float-bit buffers, for
# skeletal poses. Same conventions as fmath.ludic: float bits in words, column-
# major matrices, o may alias its inputs unless stated.
# ============================================================================
function q_new() -> words { let q = words(4); q_identity(q); return q }
function q_identity(q: words) -> void { q[0] = F_ZERO; q[1] = F_ZERO; q[2] = F_ZERO; q[3] = F_ONE }
function q_set(q: words, x: int, y: int, z: int, w: int) -> void { q[0] = x; q[1] = y; q[2] = z; q[3] = w }
function q_copy(o: words, a: words) -> void { o[0] = a[0]; o[1] = a[1]; o[2] = a[2]; o[3] = a[3] }
# the idx-th quaternion of a packed buffer
function q_load(o: words, src: words, idx: int) -> void { for i in 0 .. 4 { o[i] = src[idx * 4 + i] } }
function q_store(dst: words, idx: int, a: words) -> void { for i in 0 .. 4 { dst[idx * 4 + i] = a[i] } }
# a rotation of `angle` radians about the unit axis (ax, ay, az)
function q_axis_angle(o: words, ax: int, ay: int, az: int, angle: int) -> void {
let h = f_mul(angle, F_HALF)
let s = f_sin(h)
o[0] = f_mul(ax, s); o[1] = f_mul(ay, s); o[2] = f_mul(az, s); o[3] = f_cos(h)
}
# o = a * b (apply b first, then a); o may alias a or b
function q_mul(o: words, a: words, b: words) -> void {
let ax = a[0]; let ay = a[1]; let az = a[2]; let aw = a[3]
let bx = b[0]; let by = b[1]; let bz = b[2]; let bw = b[3]
let x = f_sub(f_add(f_add(f_mul(aw, bx), f_mul(ax, bw)), f_mul(ay, bz)), f_mul(az, by))
let y = f_add(f_add(f_sub(f_mul(aw, by), f_mul(ax, bz)), f_mul(ay, bw)), f_mul(az, bx))
let z = f_add(f_sub(f_add(f_mul(aw, bz), f_mul(ax, by)), f_mul(ay, bx)), f_mul(az, bw))
let w = f_sub(f_sub(f_sub(f_mul(aw, bw), f_mul(ax, bx)), f_mul(ay, by)), f_mul(az, bz))
o[0] = x; o[1] = y; o[2] = z; o[3] = w
}
function q_conj(o: words, a: words) -> void { o[0] = f_neg(a[0]); o[1] = f_neg(a[1]); o[2] = f_neg(a[2]); o[3] = a[3] }
function q_normalize(q: words) -> void {
let l = f_sqrt(f_add(f_add(f_mul(q[0], q[0]), f_mul(q[1], q[1])), f_add(f_mul(q[2], q[2]), f_mul(q[3], q[3]))))
if l == 0 { q_identity(q); return }
let inv = f_div(F_ONE, l)
for i in 0 .. 4 { q[i] = f_mul(q[i], inv) }
}
# normalised linear blend from a to b (shortest arc), fine for the small steps a pose takes
function q_nlerp(o: words, a: words, b: words, t: int) -> void {
var d = f_add(f_add(f_mul(a[0], b[0]), f_mul(a[1], b[1])), f_add(f_mul(a[2], b[2]), f_mul(a[3], b[3])))
var sg = F_ONE
if f_ls(d, F_ZERO) { sg = f_neg(F_ONE) }
for i in 0 .. 4 { o[i] = f_lerp(a[i], f_mul(b[i], sg), t) }
q_normalize(o)
}
# rotate the vector v by q: o = q v q*
function q_rotate(o: words, q: words, v: words) -> void {
let qx = q[0]; let qy = q[1]; let qz = q[2]; let qw = q[3]
# t = 2 * cross(q.xyz, v)
let tx = f_mul(F_TWO, f_sub(f_mul(qy, v[2]), f_mul(qz, v[1])))
let ty = f_mul(F_TWO, f_sub(f_mul(qz, v[0]), f_mul(qx, v[2])))
let tz = f_mul(F_TWO, f_sub(f_mul(qx, v[1]), f_mul(qy, v[0])))
# o = v + w t + cross(q.xyz, t)
let x = f_add(f_add(v[0], f_mul(qw, tx)), f_sub(f_mul(qy, tz), f_mul(qz, ty)))
let y = f_add(f_add(v[1], f_mul(qw, ty)), f_sub(f_mul(qz, tx), f_mul(qx, tz)))
let z = f_add(f_add(v[2], f_mul(qw, tz)), f_sub(f_mul(qx, ty), f_mul(qy, tx)))
o[0] = x; o[1] = y; o[2] = z
}
# pitch about X, yaw about Y, roll about Z, composed as yaw * pitch * roll
var q_scratch: words = null
function q_euler(o: words, pitch: int, yaw: int, roll: int) -> void {
if q_scratch == null { q_scratch = words(16) }
let qx = q_scratch; let qy = mem_off(q_scratch, 16); let qz = mem_off(q_scratch, 32); let t = mem_off(q_scratch, 48)
q_axis_angle(qx, F_ONE, F_ZERO, F_ZERO, pitch)
q_axis_angle(qy, F_ZERO, F_ONE, F_ZERO, yaw)
q_axis_angle(qz, F_ZERO, F_ZERO, F_ONE, roll)
q_mul(t, qy, qx)
q_mul(o, t, qz)
}
# the rotation matrix of q (column-major, translation cleared)
function q_to_m4(m: words, q: words) -> void {
let x = q[0]; let y = q[1]; let z = q[2]; let w = q[3]
let xx = f_mul(x, x); let yy = f_mul(y, y); let zz = f_mul(z, z)
let xy = f_mul(x, y); let xz = f_mul(x, z); let yz = f_mul(y, z)
let wx = f_mul(w, x); let wy = f_mul(w, y); let wz = f_mul(w, z)
m[0] = f_sub(F_ONE, f_mul(F_TWO, f_add(yy, zz)))
m[1] = f_mul(F_TWO, f_add(xy, wz))
m[2] = f_mul(F_TWO, f_sub(xz, wy))
m[3] = F_ZERO
m[4] = f_mul(F_TWO, f_sub(xy, wz))
m[5] = f_sub(F_ONE, f_mul(F_TWO, f_add(xx, zz)))
m[6] = f_mul(F_TWO, f_add(yz, wx))
m[7] = F_ZERO
m[8] = f_mul(F_TWO, f_add(xz, wy))
m[9] = f_mul(F_TWO, f_sub(yz, wx))
m[10] = f_sub(F_ONE, f_mul(F_TWO, f_add(xx, yy)))
m[11] = F_ZERO
m[12] = F_ZERO; m[13] = F_ZERO; m[14] = F_ZERO; m[15] = F_ONE
}
# m = translate(t) * rotate(q) * scale(s)
function m4_trs_q(m: words, tx: int, ty: int, tz: int, q: words, sx: int, sy: int, sz: int) -> void {
q_to_m4(m, q)
for r in 0 .. 3 { m[r] = f_mul(m[r], sx); m[4 + r] = f_mul(m[4 + r], sy); m[8 + r] = f_mul(m[8 + r], sz) }
m[12] = tx; m[13] = ty; m[14] = tz
}

View file

@ -0,0 +1,26 @@
# ============================================================================
# ludic.render3d — a physically based 3D renderer on Gl.* (OpenGL 4.1 core).
# Import this one file; the game supplies scene_draw() / scene_draw_casters().
# ============================================================================
import "fmath.ludic"
import "prof.ludic"
import "programs.ludic"
import "texture.ludic"
import "mesh.ludic"
import "camera.ludic"
import "sky.ludic"
import "daylight.ludic"
import "terrain.ludic"
import "collide.ludic"
import "overlay.ludic"
import "shadow.ludic"
import "post.ludic"
import "quat.ludic"
import "gltf.ludic"
import "skin.ludic"
import "scatter.ludic"
import "actor.ludic"
import "stream.ludic"
import "grass.ludic"
import "water.ludic"
import "render.ludic"

View file

@ -0,0 +1,206 @@
# ============================================================================
# render.ludic — the frame. Shadow cascades, the HDR scene pass (terrain, the
# scene's objects, the sky), bloom, and the tonemapped composite to the screen.
# The scene (what the game places in the world) hooks in through scene_draw /
# scene_draw_casters, which the demo defines.
# ============================================================================
var r3d_sky_prog: int = 0
var r3d_fog_density: int = 0
var r3d_fog_falloff: int = 0
var r3d_time: int = 0
var r3d_ready: bool = false
# set before r3d_init to build the landscape from a real height map
var r3d_dem_path: string = null
var r3d_dem_min: int = 0
var r3d_dem_max: int = 0
var r3d_dem_base: int = 0
var r3d_dem_ox: int = 0
var r3d_dem_oz: int = 0
var r3d_ortho_path: string = null
var r3d_debug: bool = false
var r3d_debug_shadow: bool = false
var r3d_debug_max: bool = false
var r3d_cloud_shadow: int = 0x3F000000 # 0.5
var r3d_clip_y: int = 0xCF000000 # -2^31: no clipping
function fog_bind(prog: int) -> void {
u_f(gl_uniform(prog, "u_clip_y"), r3d_clip_y)
u_f(gl_uniform(prog, "u_spec_scale"), F_ONE)
u_f(gl_uniform(prog, "u_fog_density"), r3d_fog_density)
u_f(gl_uniform(prog, "u_fog_height_falloff"), r3d_fog_falloff)
var cs = r3d_cloud_shadow
if Os.has_env("R3D_NOCLOUD") { cs = F_ZERO }
u_f(gl_uniform(prog, "u_cloud_shadow"), cs)
u_f(gl_uniform(prog, "u_time"), r3d_time)
}
# profiling switches (environment): R3D_NOSHADOW R3D_NOGI R3D_MSAA=n R3D_NOBLADES R3D_NOCARDS R3D_NOTREES R3D_NEAR=m
var r3d_test_frame: int = 0
var r3d_test_resize: int = 0
var r3d_no_shadow: bool = false
var r3d_no_trees: bool = false
var r3d_no_refl: bool = false
function r3d_env_flags() -> void {
r3d_no_shadow = Os.has_env("R3D_NOSHADOW")
r3d_no_trees = Os.has_env("R3D_NOTREES")
r3d_no_refl = Os.has_env("R3D_NOREFL")
if Os.has_env("R3D_DEBUG") { r3d_debug = true }
if Os.has_env("R3D_DBGSHADOW") { r3d_debug_shadow = true }
if Os.has_env("R3D_NOGI") { post_gi_strength = F_ZERO; post_ao_strength = F_ZERO; post_no_gi = true }
if Os.has_env("R3D_MSAA") { post_ms_samples = Text.to_int(Os.env("R3D_MSAA")) }
sc_skip_blade = Os.has_env("R3D_NOBLADES")
sc_skip_card = Os.has_env("R3D_NOCARDS")
sc_dbg_lod = Os.has_env("R3D_LODDBG")
if Os.has_env("R3D_ANISO") {
let a = Text.to_int(Os.env("R3D_ANISO"))
tex_anisotropy = 1.0
if a >= 2 { tex_anisotropy = 2.0 }
if a >= 4 { tex_anisotropy = 4.0 }
if a >= 8 { tex_anisotropy = 8.0 }
if a >= 16 { tex_anisotropy = 16.0 }
}
}
function r3d_init(w: int, h: int, title: string) -> bool {
r3d_env_flags()
if not gl_open(w, h, title) { print("r3d: no OpenGL context"); return false }
if Os.has_env("R3D_NOVSYNC") { gl_vsync(0) }
var renderer: string = gl_get_string(GL_RENDERER)
print(`r3d: {gl_w}x{gl_h} on {renderer}`)
prof_init()
cam_init(fr(gl_w, gl_h))
if not sky_load(r3d_assets + "/hdri/kloofendal_48d_partly_cloudy_puresky_4k.hdr") { return false }
daylight_init()
if r3d_dem_path != null { terrain_use_dem(r3d_dem_path, r3d_dem_min, r3d_dem_max, r3d_dem_base, r3d_dem_ox, r3d_dem_oz) }
if r3d_ortho_path != null { terrain_use_ortho(r3d_ortho_path) }
terrain_init()
shadow_init()
post_init(gl_w, gl_h)
scatter_init()
actor_init()
grass_init()
r3d_sky_prog = r3d_program("fullscreen.vert", "sky.frag", "")
r3d_fog_density = fl(0.00014)
r3d_fog_falloff = fl(0.002)
r3d_ready = true
gl_check("r3d init")
return true
}
function r3d_draw_sky() -> void {
gl_depth_func(GL_LEQUAL)
gl_depth_mask(0)
gl_disable(GL_CULL_FACE)
let p = r3d_sky_prog
gl_use_program(p)
r3d_bind_2d(p, "u_sky", 0, sky_tex)
sky_bind_rot(p)
sky_bind_lighting(p)
u_mat4(gl_uniform(p, "u_inv_vp"), cam_inv_vp)
u_f(gl_uniform(p, "u_sky_gain"), fl(0.95))
u_f(gl_uniform(p, "u_sky_sat"), fl(1.35))
u_f(gl_uniform(p, "u_time"), r3d_time)
mesh_draw(sky_fullscreen)
gl_depth_mask(1)
gl_depth_func(GL_LESS)
}
# the drawable changed size: the camera's aspect and every screen-sized target follow
function r3d_resize() -> void {
cam_aspect = fr(gl_w, gl_h)
cam_update()
post_free()
post_init(gl_w, gl_h)
if water_refl != null { target_free(water_refl); water_refl = null }
print(`r3d: resized to {gl_w}x{gl_h}`)
}
function r3d_frame(time: int) -> void {
if not r3d_ready { return }
if gl_resize_check() { r3d_resize() }
# R3D_RESIZE_AT=<frame>: rebuild every screen-sized buffer mid-run, as a window resize
# or a fullscreen change does. Headless has no window to resize, and this path is where
# a stale attachment or a texture freed twice shows up.
r3d_test_frame += 1
if r3d_test_resize == 0 and Os.has_env("R3D_RESIZE_AT") { r3d_test_resize = Text.to_int(Os.env("R3D_RESIZE_AT")) }
# R3D_RESIZE_AT=<n>: from frame n on, rebuild every screen-sized buffer every few
# frames at a different size, as dragging a window edge or entering fullscreen does.
if r3d_test_resize > 0 and r3d_test_frame >= r3d_test_resize and r3d_test_frame % 4 == 0 {
let step = (r3d_test_frame / 4) % 4
var w = 1920; var h = 1080
if step == 1 { w = 1440; h = 810 }
if step == 2 { w = 2560; h = 1440 }
if step == 3 { w = 1281; h = 721 }
gl_w = w; gl_h = h
r3d_resize()
}
r3d_time = time
# the frame's counters close here, before any of its own work: the window each of them
# covers is exactly one frame, from this point to the same point next time
prof_gen_frame()
prof_mark_start()
cam_begin_frame(post_frame, gl_w, gl_h)
# the height-field shadow rebakes as the light moves in steps (daylight), or with the sky yaw when there is no clock
if (not day_on and ter_shadow_yaw != sky_yaw) or ter_shadow_gen != day_gen {
ter_shadow_gen = day_gen
let t_bk = gl_now_us()
terrain_bake_shadow()
prof_bake_add(gl_now_us() - t_bk)
}
scatter_begin_frame()
prof_cpu_mark("shadow rebake")
stream_update_all()
prof_cpu_mark("streaming")
if not r3d_no_shadow { prof_begin("shadow"); shadow_pass(); prof_end() }
prof_cpu_mark("shadow pass")
if water_on and not r3d_no_refl { prof_begin("water reflection"); water_reflection_pass(); prof_end() }
prof_cpu_mark("reflection")
post_begin_scene()
prof_begin("terrain sun")
terrain_sun_prepare()
prof_end()
prof_begin("terrain")
terrain_draw()
prof_end()
prof_cpu_mark("terrain")
prof_begin("scene (vegetation)")
scene_draw()
prof_end()
prof_cpu_mark("vegetation")
prof_begin("grass")
grass_draw()
prof_end()
prof_cpu_mark("grass")
prof_begin("sky")
r3d_draw_sky()
prof_end()
prof_begin("resolve MSAA")
post_resolve()
prof_end()
# transparent water over the resolved frame: it tests against the frame's own
# depth and reads a copy of it for the depth tint and soft shores
if water_on {
prof_begin("water surface")
post_capture_scene()
target_bind(post_hdr)
gl_enable(GL_DEPTH_TEST)
gl_depth_func(GL_LESS)
water_draw(post_depth_copy.depth)
prof_end()
}
post_color = post_hdr.color
if not post_no_gi { prof_begin("SSAO/GI"); post_ssao_pass(); prof_end() }
if r3d_debug_max { tex_max(post_hdr.color, post_hdr.w, post_hdr.h, "hdr") }
prof_begin("bloom")
post_bloom_pass()
prof_end()
prof_begin("tonemap+exposure")
post_tonemap(post_color)
prof_end()
prof_cpu_mark("post")
prof_begin("prev-colour copy")
post_capture_prev()
prof_end()
prof_collect()
gl_check("frame")
}

File diff suppressed because it is too large Load diff

View file

@ -0,0 +1,23 @@
// auto-exposure on the GPU: the scene's mean luminance from the top of its mip chain,
// eased toward from the previous frame's value, written to a 1x1 texture the tonemapper
// reads. Nothing comes back to the CPU (a readback there waited for the whole frame's
// GPU work and serialised the two: 34 ms -> the sum of both, measured 2026-09-09).
in vec2 v_uv;
out vec4 o_color;
uniform sampler2D u_scene;
uniform sampler2D u_prev;
uniform float u_lod;
uniform float u_key;
uniform float u_max;
uniform float u_rate;
uniform float u_reset;
void main() {
vec3 c = textureLod(u_scene, vec2(0.5), u_lod).rgb;
c = clamp(c, vec3(0.0), vec3(1.0e5));
float lum = dot(c, vec3(0.2126, 0.7152, 0.0722));
float target = clamp(u_key / max(lum, 0.001), 0.02, u_max);
float prev = texture(u_prev, vec2(0.5)).r;
float e = mix(prev, target, u_rate);
if (u_reset > 0.5 || !(prev > 0.0)) e = target;
o_color = vec4(e, 0.0, 0.0, 1.0);
}

View file

@ -0,0 +1,29 @@
// impostor bake: albedo + coverage, and the model-frame normal
in vec3 v_wpos;
in vec3 v_nrm;
in vec2 v_uv;
in float v_seed;
layout(location = 0) out vec4 o_albedo;
layout(location = 1) out vec4 o_normal;
uniform sampler2D u_diff;
uniform sampler2D u_arm;
void main() {
vec4 d = texture(u_diff, v_uv);
if (d.a < 0.5) discard; // cut-out cards (needles, blades, leaves) bake with their shape
vec3 N = normalize(v_nrm);
if (!gl_FrontFacing) N = -N;
#ifdef FLOWER
float vy = clamp(v_uv.y, 0.0, 1.0);
if (v_uv.x >= 2.0) d.rgb = vec3(0.035, 0.09, 0.02) * (0.7 + 0.6 * vy);
else if (v_uv.x >= 1.0) {
float f = fract(v_uv.x);
vec3 violet = vec3(0.06, 0.03, 0.3);
vec3 lip = vec3(0.4, 0.3, 0.68);
// each floret: a dark keel at the base, a pale standard at the top edge
d.rgb = mix(violet, lip, smoothstep(0.35, 1.0, vy) * 0.6 + 0.25 * smoothstep(0.3, 0.0, abs(f - 0.5)));
d.rgb *= 0.85 + 0.15 * vy;
} else d.rgb = vec3(0.08, 0.17, 0.03);
#endif
o_albedo = vec4(d.rgb, 1.0);
o_normal = vec4(N * 0.5 + 0.5, texture(u_arm, v_uv).r);
}

View file

@ -0,0 +1,27 @@
// 13-tap downsample (Jimenez), with a soft threshold on the first level
in vec2 v_uv;
out vec4 o_color;
uniform sampler2D u_src;
uniform vec2 u_texel;
uniform float u_threshold; // <0: no threshold
void main() {
vec2 t = u_texel;
vec3 a = texture(u_src, v_uv + t * vec2(-2, 2)).rgb, b = texture(u_src, v_uv + t * vec2(0, 2)).rgb, c = texture(u_src, v_uv + t * vec2(2, 2)).rgb;
vec3 d = texture(u_src, v_uv + t * vec2(-2, 0)).rgb, e = texture(u_src, v_uv).rgb, f = texture(u_src, v_uv + t * vec2(2, 0)).rgb;
vec3 g = texture(u_src, v_uv + t * vec2(-2, -2)).rgb, h = texture(u_src, v_uv + t * vec2(0, -2)).rgb, i = texture(u_src, v_uv + t * vec2(2, -2)).rgb;
vec3 j = texture(u_src, v_uv + t * vec2(-1, 1)).rgb, k = texture(u_src, v_uv + t * vec2(1, 1)).rgb;
vec3 l = texture(u_src, v_uv + t * vec2(-1, -1)).rgb, m = texture(u_src, v_uv + t * vec2(1, -1)).rgb;
a = min(a, vec3(4096.0)); b = min(b, vec3(4096.0)); c = min(c, vec3(4096.0)); d = min(d, vec3(4096.0)); e = min(e, vec3(4096.0));
f = min(f, vec3(4096.0)); g = min(g, vec3(4096.0)); h = min(h, vec3(4096.0)); i = min(i, vec3(4096.0)); j = min(j, vec3(4096.0));
k = min(k, vec3(4096.0)); l = min(l, vec3(4096.0)); m = min(m, vec3(4096.0));
vec3 col = e * 0.125 + (a + c + g + i) * 0.03125 + (b + d + f + h) * 0.0625 + (j + k + l + m) * 0.125;
if (u_threshold >= 0.0) {
float br = max(col.r, max(col.g, col.b));
float knee = u_threshold * 0.5;
float soft = clamp(br - u_threshold + knee, 0.0, 2.0 * knee);
soft = soft * soft / (4.0 * knee + 1e-4);
float contrib = max(soft, br - u_threshold) / max(br, 1e-4);
col *= contrib;
}
o_color = vec4(sane(col), 1.0);
}

View file

@ -0,0 +1,13 @@
// 3x3 tent upsample, added onto the destination (blend ONE ONE)
in vec2 v_uv;
out vec4 o_color;
uniform sampler2D u_src;
uniform vec2 u_texel;
uniform float u_radius;
void main() {
vec2 t = u_texel * u_radius;
vec3 s = texture(u_src, v_uv + t * vec2(-1, 1)).rgb + texture(u_src, v_uv + t * vec2(0, 1)).rgb * 2.0 + texture(u_src, v_uv + t * vec2(1, 1)).rgb
+ texture(u_src, v_uv + t * vec2(-1, 0)).rgb * 2.0 + texture(u_src, v_uv).rgb * 4.0 + texture(u_src, v_uv + t * vec2(1, 0)).rgb * 2.0
+ texture(u_src, v_uv + t * vec2(-1, -1)).rgb + texture(u_src, v_uv + t * vec2(0, -1)).rgb * 2.0 + texture(u_src, v_uv + t * vec2(1, -1)).rgb;
o_color = vec4(s / 16.0, 1.0);
}

View file

@ -0,0 +1,7 @@
// full-screen triangle from gl_VertexID; uv in [0,1], z = 1 (the far plane)
out vec2 v_uv;
void main() {
vec2 p = vec2((gl_VertexID == 1) ? 3.0 : -1.0, (gl_VertexID == 2) ? 3.0 : -1.0);
v_uv = p * 0.5 + 0.5;
gl_Position = vec4(p, 1.0, 1.0);
}

View file

@ -0,0 +1,149 @@
// Procedural ground-cover blades, generated on the GPU, with NO distance rings.
//
// The world is cut into 16 m cells. Blade j of a cell always stands at the same place
// (a hash of the cell and j), so a blade never moves. How many of a cell's blades exist
// is a smooth function of the blade's own distance to the camera: cell area over the
// square of a spacing that grows linearly with distance. Thinning removes the highest
// indices first, and a blade shrinks before it goes, so density is continuous in space
// and in time and nothing can form a boundary. Draws are per tile (CPU-side frustum
// culling, grass.ludic); a tile only decides how many indices to feed the shader.
layout(location = 0) in vec3 a_pos; // x: -0.5..0.5 across, y: 0..1 along the blade, z: bend
layout(location = 2) in vec2 a_uv;
uniform mat4 u_view;
uniform mat4 u_proj;
uniform mat4 u_vp;
uniform vec3 u_cam_pos;
uniform float u_time;
uniform sampler2D u_ts_height;
uniform vec2 u_ts_origin;
uniform float u_ts_half;
uniform sampler2D u_ortho;
uniform float u_ortho_on;
uniform float u_lake_level;
uniform float u_snow_line;
uniform float u_wind;
uniform vec2 u_tile; // world xz of this tile's corner
uniform int u_tile_cells; // 16 m cells per tile side
uniform int u_per_cell; // indices drawn per cell in this tile
uniform float u_s0; // blade spacing at the camera (m)
uniform float u_d0; // distance at which the spacing has doubled (m)
uniform float u_radius; // no blades past this
uniform int u_dbg; // R3D_GRASS_DBG: 1 lift blades 0.3 m, 2 light as ground everywhere, 3 both
out vec3 v_wpos;
out vec3 v_nrm;
out vec2 v_uv;
out float v_seed;
out vec2 v_rot;
out float v_hull;
const float CELL = 16.0;
float hash1(vec2 p) { return fract(sin(dot(p, vec2(127.1, 311.7))) * 43758.5453123); }
// An integer hash (PCG) for the per-blade values. The sine hash advanced linearly with
// the blade index, so a cell's blades fell into diagonal rows, and the rows read as
// streaks across the meadow with an edge wherever they thinned out.
uint pcg(uint v) { uint s = v * 747796405u + 2891336453u; uint w = ((s >> ((s >> 28u) + 4u)) ^ s) * 277803737u; return (w >> 22u) ^ w; }
float bladeHash(ivec2 cell, int j, int k) {
uint h = pcg(uint(cell.x + 32768) * 73856093u ^ uint(cell.y + 32768) * 19349663u ^ uint(j) * 83492791u ^ uint(k) * 2654435761u);
return float(h) * (1.0 / 4294967295.0);
}
float heightSmooth(sampler2D tex, vec2 uv) {
vec2 res = vec2(textureSize(tex, 0));
vec2 t = uv * res - 0.5;
vec2 f = fract(t);
vec2 i = floor(t);
vec2 w0 = (1.0 - f) * (1.0 - f) * (1.0 - f) / 6.0;
vec2 w1 = (4.0 - 6.0 * f * f + 3.0 * f * f * f) / 6.0;
vec2 w3 = f * f * f / 6.0;
vec2 w2 = 1.0 - w0 - w1 - w3;
vec2 s0 = w0 + w1, s1 = w2 + w3;
vec2 o0 = (i - 1.0 + w1 / s0 + 0.5) / res;
vec2 o1 = (i + 1.0 + w3 / s1 + 0.5) / res;
return (texture(tex, vec2(o0.x, o0.y)).r * s0.x + texture(tex, vec2(o1.x, o0.y)).r * s1.x) * s0.y
+ (texture(tex, vec2(o0.x, o1.y)).r * s0.x + texture(tex, vec2(o1.x, o1.y)).r * s1.x) * s1.y;
}
void cull() { gl_Position = vec4(0.0, 0.0, 2.0, 1.0); v_wpos = vec3(0.0); v_nrm = vec3(0.0, 1.0, 0.0); v_uv = vec2(0.0); v_seed = 0.0; v_rot = vec2(0.0, 1.0); v_hull = 1.0; }
void main() {
int i = gl_InstanceID;
int c = i / u_per_cell;
int j = i - c * u_per_cell;
vec2 cell = u_tile + vec2(float(c % u_tile_cells), float(c / u_tile_cells)) * CELL;
vec2 cid = floor(cell / CELL + 0.5);
ivec2 ci = ivec2(cid);
float fj = float(j);
// this blade's fixed place in its cell
vec2 hv = vec2(bladeHash(ci, j, 0), bladeHash(ci, j, 1));
vec2 xz = cell + hv * CELL;
vec2 d2 = xz - u_cam_pos.xz;
float dist = length(d2);
if (dist >= u_radius) { cull(); return; }
// how many of this cell's blades exist at this distance: area over spacing^2, spacing
// growing linearly with distance. j beyond that count does not exist; the last fifth
// of the count shrinks to nothing so a blade never pops.
float spacing = u_s0 * (1.0 + dist / u_d0);
float count = CELL * CELL / (spacing * spacing) * (1.0 - smoothstep(u_radius * 0.7, u_radius, dist));
if (fj >= count) { cull(); return; }
float life = 1.0 - smoothstep(0.8, 1.0, fj / max(count, 1.0));
// the ground under it
vec2 huv = (xz - u_ts_origin) / (2.0 * u_ts_half) + 0.5;
if (huv.x < 0.0 || huv.x > 1.0 || huv.y < 0.0 || huv.y > 1.0) { cull(); return; }
vec4 ht = texture(u_ts_height, huv);
vec4 croot = u_vp * vec4(xz.x, ht.r, xz.y, 1.0);
if (croot.w < -1.0 || abs(croot.x) > croot.w * 1.25 + 1.5 || abs(croot.y) > croot.w * 1.4 + 1.5) { cull(); return; }
vec3 gn = normalize(ht.gba);
float h3 = bladeHash(ci, j, 2), h4 = bladeHash(ci, j, 3);
float ok = (1.0 - smoothstep(0.30, 0.55, 1.0 - gn.y)) * smoothstep(0.0, 0.6, ht.r - u_lake_level - 0.15) * smoothstep(u_snow_line - 80.0, u_snow_line - 200.0, ht.r);
if (u_ortho_on > 0.5) {
vec3 oc = textureLod(u_ortho, huv, 1.5).rgb;
ok *= 0.25 + 0.75 * smoothstep(0.0, 0.02, oc.g - max(oc.r, oc.b));
}
if (h4 > ok) { cull(); return; }
// the root on the drawn surface: the CDLOD mesh follows the B-spline to within
// centimetres near the camera, so the smooth sample is the drawn height
bool far = dist > 300.0;
float h = far ? ht.r : heightSmooth(u_ts_height, huv);
if (far) h += 0.03;
if ((u_dbg & 1) != 0) h += 0.3;
// the blade: sized so that coverage stays level as the spacing grows
float seed = hv.x * 0.7 + hv.y * 0.3;
float ang = hv.y * 6.2831853;
float s = sin(ang), c_ = cos(ang);
float grow = spacing / u_s0; // 1 at the camera, growing with distance
float tall = mix(0.18, 0.42, h3) * mix(0.8, 1.2, hash1(cid * 0.1)) * (1.0 + 0.35 * smoothstep(1.0, 12.0, grow)) * life;
float bw = 0.028 * mix(1.0, 0.45 * grow, smoothstep(1.0, 4.0, grow));
if (far) { bw = max(bw, spacing * 0.35); tall = min(tall, spacing * 0.3); }
vec3 p = vec3(a_pos.x * bw, a_pos.y * tall, a_pos.z * tall * (0.6 + 0.8 * h4));
vec3 n = vec3(0.0, 0.3, 1.0);
float gust = sin(xz.x * 0.09 + u_time * 1.1) * 0.5 + sin(xz.y * 0.13 - u_time * 0.8 + xz.x * 0.05) * 0.5;
float ph = u_time * 1.7 + seed * 6.2831 + xz.x * 0.05 + xz.y * 0.07;
float sway = (sin(ph) * 0.6 + sin(ph * 2.3 + 1.0) * 0.4 + gust) * u_wind;
float hgt = max(p.y, 0.0);
p.x += sway * hgt * hgt * 0.35;
p.z += sway * hgt * hgt * 0.15 * cos(ph * 0.7);
p = vec3(c_ * p.x + s * p.z, p.y, -s * p.x + c_ * p.z);
n = normalize(vec3(c_ * n.x + s * n.z, n.y, -s * n.x + c_ * n.z));
// stand on the ground: rotate the blade's frame from world-up to the surface normal
{
vec3 up = vec3(0.0, 1.0, 0.0);
vec3 k = cross(up, gn);
float sk = length(k), ck = gn.y;
if (sk > 1e-4) {
k /= sk;
p = p * ck + cross(k, p) * sk + k * dot(k, p) * (1.0 - ck);
n = normalize(n * ck + cross(k, n) * sk + k * dot(k, n) * (1.0 - ck));
}
}
// Blades are lit with the ground's normal from a few metres out. Lit by their own
// facing, the wind's sine field bent them in bands and the lit/unlit sides flipped in
// those bands: light and dark rows across the whole meadow. Ground-normal lighting is
// what open-world grass does (the blade's own normal only matters within arm's reach).
n = normalize(mix(n, gn, smoothstep(2.0, 12.0, dist)));
vec3 w = vec3(xz.x, h - 0.02, xz.y) + p;
v_wpos = w;
v_nrm = n;
v_uv = far ? vec2(a_uv.x, 0.45 + 0.2 * a_uv.y) : a_uv;
v_seed = seed;
v_rot = vec2(s, c_);
v_hull = (dist > 2.0 || far || (u_dbg & 2) != 0) ? -1.0 : 1.0; // no back-face flip, no rounding past arm's reach
gl_Position = u_proj * u_view * vec4(w, 1.0);
}

View file

@ -0,0 +1,110 @@
// world height (metres) for the texel's x/z; R32F target
in vec2 v_uv;
out vec4 o;
uniform float u_half;
#ifdef DEM
uniform sampler2D u_dem; // 16-bit height map of a real place (Copernicus GLO-30)
uniform float u_dem_min;
uniform float u_dem_max;
uniform float u_dem_base; // the elevation that becomes y = 0
uniform vec2 u_origin; // world x/z of the map's centre
uniform vec4 u_lake; // a lake: centre x/z, half extents (zero = none)
uniform float u_lake_level; // its surface height; the model records the surface, the bed is carved below it
// The DEM is Copernicus GLO-30 — 30 m data resampled onto this 4 m grid — so the stored
// field is piecewise linear with a slope discontinuity every ~7 texels. Differencing it
// for a shading normal turns each kink into a ridge, and on steep ground, where the same
// kink spans a large height change, they read as a regular corrugation running across
// the slope. Smoothing across that lattice removes them and discards no real detail:
// there is none below 30 m in the source, and the fractal detail below supplies the fine
// relief. Kernel is a separable gaussian sampled at 3-texel spacing (~12 m each side).
float demRaw(vec2 uv) { return texture(u_dem, uv).r; }
// u_dem_blur texels of separable gaussian (0 = the survey as it is). The 30 m Copernicus
// model needed ~3 texels to hide its resampling lattice; 2 m lidar needs none.
uniform float u_dem_blur;
float demH(vec2 uv) {
if (u_dem_blur <= 0.0) return mix(u_dem_min, u_dem_max, demRaw(uv));
vec2 t = u_dem_blur / vec2(textureSize(u_dem, 0));
float c = demRaw(uv);
float e = demRaw(uv + vec2(t.x, 0.0)) + demRaw(uv - vec2(t.x, 0.0))
+ demRaw(uv + vec2(0.0, t.y)) + demRaw(uv - vec2(0.0, t.y));
float d = demRaw(uv + t) + demRaw(uv - t)
+ demRaw(uv + vec2(t.x, -t.y)) + demRaw(uv + vec2(-t.x, t.y));
float f = demRaw(uv + 2.0 * vec2(t.x, 0.0)) + demRaw(uv - 2.0 * vec2(t.x, 0.0))
+ demRaw(uv + 2.0 * vec2(0.0, t.y)) + demRaw(uv - 2.0 * vec2(0.0, t.y));
float h = (12.0 * c + 6.0 * e + 3.0 * d + 2.0 * f) / (12.0 + 24.0 + 12.0 + 8.0);
return mix(u_dem_min, u_dem_max, h);
}
#endif
float bump(vec2 p, vec2 c, float r) { float d = length(p - c) / r; return exp(-d * d * 2.0); }
void main() {
#ifdef DEM
vec2 xz = (v_uv - 0.5) * 2.0 * u_half + u_origin;
float e = 1.0 / 2048.0;
float h = demH(v_uv) - u_dem_base;
// the basin mask keys off the UNSMOOTHED sample: the lake is only ~1 m below its shore
// in the model, and the 12 m blur above averages the shore with the flat water beside
// it, dragging it under the outline threshold — the camera's own bank was carved 7 m down
float raw = demRaw(v_uv) * (u_dem_max - u_dem_min) + u_dem_min - u_dem_base;
// the survey carries its own relief at this resolution; only the foot track is added
h -= 0.25 * smoothstep(6.0, 1.5, pathDist(xz));
if (u_lake.z > 0.0) {
// the elevation model samples the water surface as flat ground: inside the lake's
// outline sink it into a bed, deepest in the middle, with a gentle gravel ramp at the shore
vec2 q = (xz - u_lake.xy) / u_lake.zw;
float inside = smoothstep(1.0, 0.8, dot(q, q));
// the lidar records the lake as a flat surface exactly at its level: everything at or
// just above that level inside the outline is lake bed
float basin = smoothstep(u_lake_level + 0.6, u_lake_level - 0.3, raw) * inside;
float bed = u_lake_level - 0.4 - (5.0 + 3.0 * (1.0 - dot(q, q)) + 1.5 * fbm(xz * 0.02, 3)) * basin * basin; // a gentle gravel ramp, then the drop
h = mix(h, bed, basin);
}
o = vec4(h, 0.0, 0.0, 1.0);
#elif defined(SMOOTH)
// A purpose-built test ground: 2 km square, analytically smooth everywhere. No survey
// data, so no resampling lattice and no quantised source — if a grid still shows on
// this, the cause is in the renderer rather than the elevation model.
vec2 xz = (v_uv - 0.5) * 2.0 * u_half;
float d = length(xz);
// meadow: a gentle roll a few metres either side of 8 m, two octaves only
float h = 8.0 + 4.5 * fbm(xz * 0.0018, 3) + 1.4 * fbm(xz * 0.007, 3);
// a small pond in the middle: a smooth basin ~90 m across, floor about 5 m down
float pond = exp(-dot(xz, xz) / (2.0 * 70.0 * 70.0));
h -= 13.0 * pond;
// the rim mountain, 200 m above the meadow, starting well outside the grass
h += smoothstep(620.0, 980.0, d) * (200.0 + 90.0 * fbm(xz * 0.0025 + 4.0, 4));
o = vec4(h, 0.0, 0.0, 1.0);
#else
vec2 xz = (v_uv - 0.5) * 2.0 * u_half;
float d = length(xz);
// the meadow falls away northward (-z) from the rise the camera stands on, into a broad valley
float base = 0.055 * xz.y + 4.0 * fbm(xz * 0.008, 5) + 1.2 * fbm(xz * 0.04, 4) + 0.15 * fbm(xz * 0.35, 3);
// soft valley floor with a stream
float floorY = -22.0;
float k = 12.0;
base = floorY + k * log(1.0 + exp((base - floorY) / k));
float sb = smoothstep(6.0, 0.0, abs(xz.x - 150.0 - 90.0 * sin(xz.y * 0.006 + 1.0))) * smoothstep(-220.0, -400.0, xz.y);
base -= 1.6 * sb;
// the track sits slightly worn in
base -= 0.25 * smoothstep(6.0, 1.5, pathDist(xz));
float h = base;
// far ridge across the valley, forested hills
float back = smoothstep(-550.0, -1000.0, xz.y);
h += back * (90.0 + 160.0 * ridged(xz * 0.0025 + 5.0, 6));
// the great snow mountain to the north-west, and its shoulder
float m1 = bump(xz, vec2(-820.0, -520.0), 520.0);
float m2 = bump(xz, vec2(-1150.0, -900.0), 600.0);
float m3 = bump(xz, vec2(-560.0, -150.0), 260.0);
float mtn = max(m1 * 620.0, max(m2 * 720.0, m3 * 180.0));
h += mtn * (0.45 + 0.55 * ridged(xz * 0.0018 + 11.0, 7)) + 40.0 * (m1 + m2) * ridged(xz * 0.008 + 3.0, 5);
// hills to the east, lower and rounder
float e1 = bump(xz, vec2(760.0, -420.0), 420.0);
float e2 = bump(xz, vec2(950.0, 150.0), 380.0);
h += (e1 * 170.0 + e2 * 120.0) * (0.6 + 0.4 * fbm(xz * 0.004 + 9.0, 5));
// the outer rim so nothing ends at a flat edge
h += smoothstep(750.0, 1024.0, d) * (120.0 + 120.0 * ridged(xz * 0.003, 5));
// a rocky knoll on the right of the meadow
float knoll = 14.0 * bump(xz, vec2(230.0, -40.0), 70.0);
h += knoll * (0.6 + 0.4 * fbm(xz * 0.05, 4));
o = vec4(h, 0.0, 0.0, 1.0);
#endif
}

View file

@ -0,0 +1,36 @@
in vec2 v_uv;
out vec4 o_color;
vec2 hammersley(uint i, uint n) {
uint b = i;
b = (b << 16u) | (b >> 16u);
b = ((b & 0x55555555u) << 1u) | ((b & 0xAAAAAAAAu) >> 1u);
b = ((b & 0x33333333u) << 2u) | ((b & 0xCCCCCCCCu) >> 2u);
b = ((b & 0x0F0F0F0Fu) << 4u) | ((b & 0xF0F0F0F0u) >> 4u);
b = ((b & 0x00FF00FFu) << 8u) | ((b & 0xFF00FF00u) >> 8u);
return vec2(float(i) / float(n), float(b) * 2.3283064365386963e-10);
}
void main() {
float NoV = max(v_uv.x, 1e-3);
float rough = max(v_uv.y, 0.02);
vec3 v = vec3(sqrt(1.0 - NoV * NoV), 0.0, NoV);
float a = rough * rough;
float A = 0.0, B = 0.0;
const uint N = 512u;
for (uint i = 0u; i < N; i++) {
vec2 x = hammersley(i, N);
float phi = 2.0 * PI * x.x;
float ct = sqrt((1.0 - x.y) / (1.0 + (a * a - 1.0) * x.y));
float st = sqrt(1.0 - ct * ct);
vec3 h = vec3(cos(phi) * st, sin(phi) * st, ct);
vec3 l = 2.0 * dot(v, h) * h - v;
float NoL = max(l.z, 0.0), NoH = max(h.z, 0.0), VoH = max(dot(v, h), 0.0);
if (NoL > 0.0) {
float G = V_Smith(NoV, NoL, a) * 4.0 * NoL * NoV; // Smith G from the visibility term
float Gv = G * VoH / max(NoH * NoV, 1e-4);
float Fc = pow(1.0 - VoH, 5.0);
A += (1.0 - Fc) * Gv;
B += Fc * Gv;
}
}
o_color = vec4(clamp(A / float(N), 0.0, 1.0), clamp(B / float(N), 0.0, 1.0), 0.0, 1.0);
}

View file

@ -0,0 +1,36 @@
in vec2 v_uv;
out vec4 o_color;
uniform sampler2D u_sky;
uniform float u_sun_clip; // clamp the sun's radiance so it does not alias the convolution
vec3 dirFromUV(vec2 uv) {
float phi = (uv.x - 0.5) * 2.0 * PI;
float theta = uv.y * PI;
return vec3(sin(theta) * sin(phi), cos(theta), -sin(theta) * cos(phi));
}
vec2 hammersley(uint i, uint n) {
uint b = i;
b = (b << 16u) | (b >> 16u);
b = ((b & 0x55555555u) << 1u) | ((b & 0xAAAAAAAAu) >> 1u);
b = ((b & 0x33333333u) << 2u) | ((b & 0xCCCCCCCCu) >> 2u);
b = ((b & 0x0F0F0F0Fu) << 4u) | ((b & 0xF0F0F0F0u) >> 4u);
b = ((b & 0x00FF00FFu) << 8u) | ((b & 0xFF00FF00u) >> 8u);
return vec2(float(i) / float(n), float(b) * 2.3283064365386963e-10);
}
void main() {
vec3 n = dirFromUV(v_uv);
vec3 up = abs(n.y) < 0.999 ? vec3(0, 1, 0) : vec3(1, 0, 0);
vec3 t = normalize(cross(up, n));
vec3 b = cross(n, t);
vec3 acc = vec3(0.0);
const uint N = 512u;
for (uint i = 0u; i < N; i++) {
vec2 h = hammersley(i, N);
float phi = 2.0 * PI * h.x;
float ct = sqrt(1.0 - h.y); // cosine-weighted
float st = sqrt(h.y);
vec3 d = t * (cos(phi) * st) + b * (sin(phi) * st) + n * ct;
vec3 c = textureLod(u_sky, skyUV(d), 5.0).rgb;
acc += min(c, vec3(u_sun_clip));
}
o_color = vec4(acc / float(N), 1.0);
}

View file

@ -0,0 +1,53 @@
in vec2 v_uv;
out vec4 o_color;
uniform sampler2D u_sky;
uniform float u_rough;
uniform float u_sun_clip;
uniform float u_sky_w;
vec3 dirFromUV(vec2 uv) {
float phi = (uv.x - 0.5) * 2.0 * PI;
float theta = uv.y * PI;
return vec3(sin(theta) * sin(phi), cos(theta), -sin(theta) * cos(phi));
}
vec2 hammersley(uint i, uint n) {
uint b = i;
b = (b << 16u) | (b >> 16u);
b = ((b & 0x55555555u) << 1u) | ((b & 0xAAAAAAAAu) >> 1u);
b = ((b & 0x33333333u) << 2u) | ((b & 0xCCCCCCCCu) >> 2u);
b = ((b & 0x0F0F0F0Fu) << 4u) | ((b & 0xF0F0F0F0u) >> 4u);
b = ((b & 0x00FF00FFu) << 8u) | ((b & 0xFF00FF00u) >> 8u);
return vec2(float(i) / float(n), float(b) * 2.3283064365386963e-10);
}
void main() {
vec3 n = dirFromUV(v_uv);
vec3 v = n;
if (u_rough < 0.02) { o_color = vec4(textureLod(u_sky, skyUV(n), 0.0).rgb, 1.0); return; }
vec3 up = abs(n.y) < 0.999 ? vec3(0, 1, 0) : vec3(1, 0, 0);
vec3 t = normalize(cross(up, n));
vec3 b = cross(n, t);
float a = u_rough * u_rough;
vec3 acc = vec3(0.0);
float wsum = 0.0;
const uint N = 256u;
for (uint i = 0u; i < N; i++) {
vec2 x = hammersley(i, N);
float phi = 2.0 * PI * x.x;
float ct = sqrt((1.0 - x.y) / (1.0 + (a * a - 1.0) * x.y));
float st = sqrt(1.0 - ct * ct);
vec3 h = t * (cos(phi) * st) + b * (sin(phi) * st) + n * ct;
vec3 l = 2.0 * dot(v, h) * h - v;
float NoL = dot(n, l);
if (NoL > 0.0) {
float NoH = max(ct, 0.0);
float D = D_GGX(NoH, a);
float pdf = D * NoH / (4.0 * max(dot(v, h), 1e-4)) + 1e-4;
float saTexel = 4.0 * PI / (u_sky_w * u_sky_w * 0.5);
float saSample = 1.0 / (float(N) * pdf);
float mip = clamp(0.5 * log2(saSample / saTexel) + 1.0, 0.0, 8.0);
vec3 c = textureLod(u_sky, skyUV(l), mip).rgb;
acc += min(c, vec3(u_sun_clip)) * NoL;
wsum += NoL;
}
}
o_color = vec4(acc / max(wsum, 1e-4), 1.0);
}

View file

@ -0,0 +1,66 @@
in vec2 v_uv;
in vec3 v_wpos;
in float v_seed;
in float v_tile;
in float v_yaw;
out vec4 o_color;
uniform sampler2D u_atlas_albedo;
uniform sampler2D u_atlas_normal;
uniform mat4 u_view;
uniform float u_tiles;
uniform vec3 u_tint;
uniform float u_radius;
void main() {
if (v_wpos.y < u_clip_y) discard;
vec2 uv = vec2((v_tile + v_uv.x) / u_tiles, v_uv.y);
#ifdef SHADOW_PASS
// the map's texels are coarse: read a finer mip so the crown's coverage is not averaged away
vec4 a = texture(u_atlas_albedo, uv, -3.0);
if (a.a < 0.22) discard;
#else
vec4 a = texture(u_atlas_albedo, uv);
float rawA = a.a;
a.rgb /= max(a.a, 1e-3); // the atlas mips are premultiplied by coverage
// alpha-to-coverage, sharpened per mip so the silhouette stays crisp at any distance
float cov = (a.a - 0.3) / max(fwidth(a.a), 1e-4) + 0.5;
if (cov < 0.02) discard;
a.a = clamp(cov, 0.0, 1.0);
#endif
#ifdef SHADOW_PASS
return;
#else
// The normal atlas is cleared to black where nothing was drawn, and its mips average
// that black into every silhouette texel — decoded, a half-covered texel pointed away
// from everything and shaded near black, so each tree wore a dark outline. Dividing by
// the same coverage restores the normal (and the AO in .a) of the covered part.
vec4 nn = texture(u_atlas_normal, uv) / max(rawA, 1e-3);
vec3 n = clamp(nn.rgb, 0.0, 1.0) * 2.0 - 1.0;
// the tile was baked from angle tile*2pi/tiles around the canonical tree; rotate by the instance yaw
float s = sin(v_yaw), c = cos(v_yaw);
n = vec3(c * n.x + s * n.z, n.y, -s * n.x + c * n.z);
n = normalize(n);
// canopy hull normal: a rounded shell over the card, blended with the baked detail
float far = smoothstep(200.0, 1200.0, length(v_wpos - u_cam_pos));
vec2 q = vec2(v_uv.x * 2.0 - 1.0, v_uv.y * 2.0 - 1.0);
vec3 toCam = normalize(u_cam_pos - v_wpos); toCam.y = 0.0; toCam = normalize(toCam);
vec3 right = vec3(-toCam.z, 0.0, toCam.x);
vec3 hull = normalize(right * q.x * 0.8 + vec3(0.0, 1.0, 0.0) * (q.y * 0.6 + 0.35) + toCam * 0.7);
n = normalize(mix(hull, n, mix(0.65, 0.35, far)));
float dist = length(v_wpos - u_cam_pos);
float viewDepth = -(u_view * vec4(v_wpos, 1.0)).z;
vec3 alb = a.rgb * u_tint * (0.85 + 0.3 * fract(v_seed * 7.13)) * regionTint(v_wpos, 0.4);
// a distant stand reads as a dark mass, not as bright separate sprites
alb = mix(alb, alb * vec3(0.72, 0.78, 0.72), far);
// the card itself is the caster: look up the shadow a little toward the sun so it does not self-shadow
float shadow = sunShadow(v_wpos + u_sun_dir * u_radius * 0.7, vec3(0, 1, 0), viewDepth);
// crowns are dense: darken toward the centre of the card as a cheap interior occlusion
float interior = 1.0 - 0.45 * smoothstep(0.9, 0.3, abs(q.x)) * smoothstep(1.0, 0.2, v_uv.y);
// ground contact: the lowest part of anything sitting on the ground is occluded by it
// (a boulder's underside, a trunk's base); without it a far boulder is a sticker on the grass
interior *= mix(0.55, 1.0, smoothstep(0.0, 0.3, v_uv.y));
vec3 col = shade(v_wpos, n, alb, 0.85, 0.0, clamp(nn.a, 0.0, 1.0) * 0.8 * interior, shadow * interior, viewDepth);
col += alb * skyIrradiance(vec3(0, 1, 0)) * 0.06;
col = applyFog(col, v_wpos, dist);
o_color = vec4(sane(col), a.a);
#endif
}

View file

@ -0,0 +1,42 @@
// camera-facing (around y) card per instance, showing the atlas tile nearest the view angle
layout(location = 0) in vec2 a_xy; // [-0.5, 0.5] x [0, 1]
layout(location = 1) in vec2 a_uv;
layout(location = 3) in vec4 i_pos; // x y z scale
layout(location = 4) in vec4 i_rot; // sin cos seed wind
uniform mat4 u_view;
uniform mat4 u_proj;
uniform mat4 u_light_vp;
uniform vec3 u_cam_pos;
uniform vec3 u_face_dir; // direction the cards face (to the camera, or the sun in the shadow pass)
uniform float u_radius;
uniform float u_height;
uniform float u_tiles;
out vec2 v_uv;
out vec3 v_wpos;
out float v_seed;
out float v_tile;
out float v_yaw;
void main() {
vec3 center = i_pos.xyz;
#ifdef SHADOW_PASS
vec3 toCam = normalize(vec3(u_face_dir.x, 0.0, u_face_dir.z));
#else
vec3 toCam = u_cam_pos - center; toCam.y = 0.0; toCam = normalize(toCam);
#endif
vec3 right = vec3(-toCam.z, 0.0, toCam.x);
float yaw = atan(i_rot.x, i_rot.y);
// angle of the viewer around the (rotated) tree, in tile units
float ang = atan(toCam.x, -toCam.z) - yaw;
float t = floor(fract(ang / 6.2831853) * u_tiles + 0.5);
v_tile = mod(t, u_tiles);
v_yaw = yaw;
vec3 w = center + right * (a_xy.x * 2.0 * u_radius * i_pos.w) + vec3(0.0, a_xy.y * u_height * i_pos.w, 0.0);
v_wpos = w;
v_uv = a_uv;
v_seed = i_rot.z;
#ifdef SHADOW_PASS
gl_Position = u_light_vp * vec4(w, 1.0);
#else
gl_Position = u_proj * u_view * vec4(w, 1.0);
#endif
}

View file

@ -0,0 +1,272 @@
// ---- PBR + IBL + cascaded shadows + aerial perspective (shared) ---------------------
uniform sampler2D u_irradiance; // equirect, diffuse-convolved sky
uniform sampler2DArray u_prefilter; // equirect, GGX-prefiltered sky per roughness level
uniform sampler2D u_brdf; // split-sum BRDF LUT
#define CASCADES 5
uniform sampler2DArrayShadow u_shadow; // CASCADES layers
float shadowTap(vec2 uv, int c, float ref) { return texture(u_shadow, vec4(uv, float(c), ref)); }
uniform mat4 u_cascade_vp[CASCADES];
uniform float u_cascade_split[CASCADES]; // view-space far distance of each cascade
uniform float u_cascade_range[CASCADES]; // light-frustum depth extent of each cascade (m)
uniform float u_cascade_texel[CASCADES]; // shadow texel size of each cascade (m)
uniform vec3 u_sun_dir; // toward the sun
uniform vec3 u_sun_color; // radiance
uniform vec3 u_cam_pos;
uniform float u_prefilter_levels;
uniform float u_fog_density;
uniform float u_fog_height_falloff;
uniform float u_clip_y; // planar-reflection pass: discard everything below this height
uniform float u_spec_scale; // 1 for surfaces; foliage crowns get a fraction: needles are
// tiny rough cylinders, not sheets, and a crown of card quads
// seen at grazing angles otherwise mirrors the sky and frosts
const float PI = 3.14159265359;
// never let a NaN or an infinity reach the frame: it would smear through the bloom pyramid
vec3 sane(vec3 c) { return (any(isnan(c)) || any(isinf(c))) ? vec3(0.0) : c; }
vec2 equirectUV(vec3 d) {
return vec2(atan(d.x, -d.z) / (2.0 * PI) + 0.5, acos(clamp(d.y, -1.0, 1.0)) / PI);
}
// the HDRI itself is read through a yaw rotation (u_sky_rot = sin, cos), so the sun can be
// placed where the scene wants it; the convolved maps are built through the same rotation
uniform vec2 u_sky_rot;
vec2 skyUV(vec3 d) {
vec3 r = vec3(u_sky_rot.y * d.x + u_sky_rot.x * d.z, d.y, -u_sky_rot.x * d.x + u_sky_rot.y * d.z);
return equirectUV(r);
}
// The HDRI is a pure sky: below the horizon it is a flat bright grey, not ground. Anything
// whose normal points down — the underside of a needle card, the lower half of a crown —
// was lighting itself from that grey and came out white. Below the horizon the light is
// what the ground reflects: the horizon sky times a meadow albedo.
const vec3 GROUND_ALB = vec3(0.30, 0.34, 0.14);
// the time of day (daylight.ludic): the sky's light scaled toward night, and the campfire
uniform vec3 u_ibl_scale;
uniform float u_daylight;
uniform vec3 u_fire_pos;
uniform vec3 u_fire_color;
uniform vec3 u_hand_pos; // a torch or flashlight in the hand
uniform vec3 u_hand_color;
uniform vec3 u_hand_dir;
uniform float u_hand_cone; // cos of the half-angle; <= -1: a point light
vec3 skyIrradianceRaw(vec3 n) { return texture(u_irradiance, equirectUV(n)).rgb * u_ibl_scale; }
vec3 skyIrradiance(vec3 n) {
vec3 up = skyIrradianceRaw(vec3(n.x, max(n.y, 0.0), n.z));
vec3 ground = skyIrradianceRaw(normalize(vec3(n.x, 0.15, n.z) + vec3(1e-4, 0.0, 0.0))) * GROUND_ALB;
return mix(ground, up, smoothstep(-0.25, 0.2, n.y));
}
vec3 skyPrefilteredRaw(vec3 r, float rough) {
float lv = rough * (u_prefilter_levels - 1.0);
float l0 = floor(lv);
float l1 = min(l0 + 1.0, u_prefilter_levels - 1.0);
vec2 uv = equirectUV(r);
return mix(texture(u_prefilter, vec3(uv, l0)).rgb, texture(u_prefilter, vec3(uv, l1)).rgb, lv - l0) * u_ibl_scale;
}
// the campfire: one warm point light, out by twelve metres
vec3 fireLight(vec3 wpos, vec3 n, vec3 albedo) {
vec3 d = u_fire_pos - wpos;
float r2 = max(dot(d, d), 0.04);
vec3 l = d * inversesqrt(r2);
float att = smoothstep(14.0, 5.0, sqrt(r2)) / (0.6 + r2);
return albedo / PI * u_fire_color * max(dot(n, l), 0.0) * att;
}
vec3 handLight(vec3 wpos, vec3 n, vec3 albedo) {
vec3 d = u_hand_pos - wpos;
float r2 = max(dot(d, d), 0.04);
vec3 l = d * inversesqrt(r2);
float att = smoothstep(26.0, 6.0, sqrt(r2)) / (0.5 + r2 * 0.35);
if (u_hand_cone > -1.0) {
float c = dot(-l, u_hand_dir);
att *= smoothstep(u_hand_cone, u_hand_cone + 0.12, c);
}
return albedo / PI * u_hand_color * max(dot(n, l), 0.0) * att;
}
vec3 skyPrefiltered(vec3 r, float rough) {
vec3 up = skyPrefilteredRaw(vec3(r.x, max(r.y, 0.0), r.z), rough);
vec3 ground = skyPrefilteredRaw(normalize(vec3(r.x, 0.15, r.z) + vec3(1e-4, 0.0, 0.0)), max(rough, 0.6)) * GROUND_ALB;
return mix(ground, up, smoothstep(-0.2, 0.15, r.y));
}
float D_GGX(float NoH, float a) { float a2 = a * a; float d = NoH * NoH * (a2 - 1.0) + 1.0; return a2 / (PI * d * d); }
float V_Smith(float NoV, float NoL, float a) {
float a2 = a * a;
float gv = NoL * sqrt(NoV * NoV * (1.0 - a2) + a2);
float gl = NoV * sqrt(NoL * NoL * (1.0 - a2) + a2);
return 0.5 / max(gv + gl, 1e-4);
}
vec3 F_Schlick(float VoH, vec3 f0) { float f = pow(1.0 - VoH, 5.0); return f0 + (1.0 - f0) * f; }
vec3 F_SchlickRough(float NoV, vec3 f0, float rough) { return f0 + (max(vec3(1.0 - rough), f0) - f0) * pow(1.0 - NoV, 5.0); }
// interleaved-gradient noise for rotated PCF taps
float ign(vec2 p) { return fract(52.9829189 * fract(0.06711056 * p.x + 0.00583715 * p.y)); }
float cascadeRange(int c) { return u_cascade_range[c]; }
float cascadeTexel(int c) { return u_cascade_texel[c]; }
// biasWorld in metres; the receiver is pushed along its normal by a texel first
float shadowSample(int c, vec3 wpos, float biasWorld) {
vec4 lp = u_cascade_vp[c] * vec4(wpos, 1.0);
vec3 p = lp.xyz / lp.w * 0.5 + 0.5;
if (p.x < 0.0 || p.x > 1.0 || p.y < 0.0 || p.y > 1.0 || p.z > 1.0) return 1.0;
float bias = biasWorld / cascadeRange(c);
float texel = 1.0 / 2048.0;
float r = ign(gl_FragCoord.xy) * 6.2831853;
float cs = cos(r), sn = sin(r);
mat2 rot = mat2(cs, sn, -sn, cs);
float s = 0.0;
const vec2 taps[8] = vec2[8](vec2(-0.7071, 0.7071), vec2(-0.0, -0.875), vec2(0.5303, 0.5303), vec2(-0.625, -0.0),
vec2(0.3536, -0.3536), vec2(-0.0, 0.375), vec2(-0.1768, -0.1768), vec2(0.125, 0.0));
// the far cascades' texels are metres wide: a wider filter turns their staircase into a penumbra
float rad = texel * ((c >= 4) ? 2.6 : (c == 3) ? 2.0 : 1.5);
for (int i = 0; i < 8; i++) {
vec2 off = rot * taps[i] * rad;
s += shadowTap(p.xy + off, c, p.z - bias);
}
return s / 8.0;
}
// one cascade's lookup, with a normal offset and a slope-scaled depth bias
float shadowSlope(int c, vec3 wpos, vec3 n, float tanT) {
float tx = cascadeTexel(c);
float filt = (c >= 4) ? 2.6 : ((c == 3) ? 2.0 : 1.5); // matches shadowSample's rad
vec3 wp = wpos + n * tx * (2.5 + 1.5 * tanT);
return shadowSample(c, wp, tx * (1.0 + filt * tanT) + 0.02);
}
// The baked height-field shadow (tershadow.frag): R = the lowest lit height over this
// ground texel, G = distance to the occluder that set it. Any receiver — ground, crown,
// card, water — compares its own height, so everything agrees on where the hill's
// shadow falls. The penumbra widens with the occluder's distance like a real one.
uniform sampler2D u_tershadow;
uniform vec2 u_ts_origin;
uniform float u_ts_half;
uniform float u_ts_on;
uniform sampler2D u_ts_height;
// the ground's normal under a world position (4 m texels): cover standing on the ground
// is lit with this beyond a few tens of metres, so a hillside and the grass on it agree
vec3 terrainNormalAt(vec3 wpos) {
vec2 uv = (wpos.xz - u_ts_origin) / (2.0 * u_ts_half) + 0.5;
float step = 1.0 / float(textureSize(u_ts_height, 0).x);
float world = step * 2.0 * u_ts_half;
float hl = texture(u_ts_height, uv - vec2(step, 0)).r, hr = texture(u_ts_height, uv + vec2(step, 0)).r;
float hd = texture(u_ts_height, uv - vec2(0, step)).r, hu = texture(u_ts_height, uv + vec2(0, step)).r;
return normalize(vec3(hl - hr, 2.0 * world, hd - hu));
}
float terrainShadow(vec3 wpos) {
if (u_ts_on < 0.5) return 1.0;
vec2 uv = (wpos.xz - u_ts_origin) / (2.0 * u_ts_half) + 0.5;
if (uv.x < 0.0 || uv.x > 1.0 || uv.y < 0.0 || uv.y > 1.0) return 1.0;
vec2 s = texture(u_tershadow, uv).rg;
float w = 0.6 + 0.02 * s.y;
return smoothstep(-w, w, wpos.y + 0.25 - s.x);
}
uniform int u_force_cascade;
// the far-field version: one hardware 2x2 tap in the cascade, no rotated disc, no blend
float sunShadowCheap(vec3 wpos, vec3 n, float viewDepth) {
int c = CASCADES - 1;
for (int i = 0; i < CASCADES - 1; i++) { if (viewDepth < u_cascade_split[i]) { c = i; break; } }
float tx = cascadeTexel(c);
vec4 lp = u_cascade_vp[c] * vec4(wpos + n * tx * 2.5, 1.0);
vec3 p = lp.xyz / lp.w * 0.5 + 0.5;
float s = 1.0;
if (p.x >= 0.0 && p.x <= 1.0 && p.y >= 0.0 && p.y <= 1.0 && p.z <= 1.0) s = shadowTap(p.xy, c, p.z - (tx * 2.0 + 0.02) / cascadeRange(c));
return min(s, terrainShadow(wpos));
}
float sunShadow(vec3 wpos, vec3 n, float viewDepth) {
int c = CASCADES - 1;
if (u_force_cascade >= 0) { float tx0 = cascadeTexel(u_force_cascade); return shadowSample(u_force_cascade, wpos + n * tx0 * 1.5, tx0 * 1.5 + 0.02); }
for (int i = 0; i < CASCADES - 1; i++) { if (viewDepth < u_cascade_split[i]) { c = i; break; } }
float NoL = max(dot(n, u_sun_dir), 0.0);
// Depth across one shadow texel changes by texel * tan(theta) on a surface lit at
// theta from its normal, and the PCF disc reaches `filt` texels out, so the bias must
// cover the drop over the whole filter rather than a single texel. The old form used
// (1 - NoL): at NoL = 0.2 that is 0.8 where tan(theta) is 4.9, six times short. With
// the caster and receiver now the same mesh, that shortfall is what let the terrain
// shadow itself along its own triangle edges — a faint grid over the whole slope.
float tanT = min(sqrt(max(1.0 - NoL * NoL, 0.0)) / max(NoL, 0.05), 10.0);
float s = shadowSlope(c, wpos, n, tanT);
// blend across the cascade edge
float edge = u_cascade_split[c];
float f = smoothstep(edge * 0.85, edge, viewDepth);
if (f > 0.0 && c < CASCADES - 1) {
s = mix(s, shadowSlope(c + 1, wpos, n, tanT), f);
}
return min(s, terrainShadow(wpos));
}
// patchy sunlight: a cloud layer projected along the sun onto the ground
uniform float u_cloud_shadow; // strength
uniform float u_time;
// the mask is baked into the height-field shadow texture's B (tershadow.frag); the
// projection along the sun and the drift are a uv shift
float cloudShadow(vec3 wpos) {
if (u_cloud_shadow <= 0.0 || u_ts_on < 0.5) return 1.0;
float h = 1400.0 - wpos.y;
vec2 c = wpos.xz + u_sun_dir.xz / max(u_sun_dir.y, 0.1) * h;
c += vec2(u_time * 3.0, u_time * 1.2);
vec2 uv = (c - u_ts_origin) / (2.0 * u_ts_half) + 0.5;
if (uv.x < 0.0 || uv.x > 1.0 || uv.y < 0.0 || uv.y > 1.0) return 1.0;
return 1.0 - u_cloud_shadow * texture(u_tershadow, uv).b;
}
// regional vegetation colour: aspen groves and drier ridges read lighter and yellower than
// the dark spruce and the lush hollows (a slow noise over the world, shaped by elevation)
vec3 regionTint(vec3 wpos, float strength) {
float n = fbm(wpos.xz * 0.0018 + 4.0, 3) * 0.5 + 0.5;
float aspen = smoothstep(0.52, 0.7, n) * smoothstep(520.0, 250.0, wpos.y);
float dry = smoothstep(0.35, 0.15, fbm(wpos.xz * 0.004 + 9.0, 3) * 0.5 + 0.5);
vec3 t = vec3(1.0);
t = mix(t, vec3(1.25, 1.3, 0.85), aspen * strength);
t = mix(t, vec3(1.15, 1.05, 0.7), dry * strength * 0.6);
return t;
}
// direct + image-based lighting for one surface
vec3 shade(vec3 wpos, vec3 n, vec3 albedo, float rough, float metal, float ao, float shadow, float viewDepth) {
vec3 v = normalize(u_cam_pos - wpos);
vec3 l = u_sun_dir;
vec3 h = normalize(v + l);
float NoV = max(dot(n, v), 1e-2);
float NoL = max(dot(n, l), 0.0);
float NoH = max(dot(n, h), 0.0);
float VoH = max(dot(v, h), 0.0);
rough = clamp(rough, 0.045, 1.0);
float a = rough * rough;
vec3 f0 = mix(vec3(0.04), albedo, metal);
vec3 F = F_Schlick(VoH, f0);
vec3 spec = min(D_GGX(NoH, a) * V_Smith(NoV, NoL, a), 12.0) * F * u_spec_scale; // cap the highlight: no half-float overflow, no fireflies
vec3 kd = (1.0 - F) * (1.0 - metal);
vec3 direct = (kd * albedo / PI + spec) * u_sun_color * NoL * shadow * cloudShadow(wpos);
// IBL
vec3 Fr = F_SchlickRough(NoV, f0, rough);
vec3 kdi = (1.0 - Fr) * (1.0 - metal);
vec3 irr = skyIrradiance(n);
vec3 diffuseIBL = kdi * albedo * irr;
vec3 r = reflect(-v, n);
vec3 pre = skyPrefiltered(r, rough);
vec2 brdf = texture(u_brdf, vec2(NoV, rough)).rg;
vec3 specIBL = pre * (Fr * brdf.x + brdf.y) * u_spec_scale;
// one bounce off the sunlit ground onto whatever faces it (a warm fill from below)
vec3 groundAlb = vec3(0.16, 0.2, 0.07);
vec3 bounce = kdi * albedo * groundAlb * (u_sun_color * max(u_sun_dir.y, 0.0) / PI + irr) * clamp(0.5 - 0.5 * n.y, 0.0, 1.0) * 0.5;
// specular occlusion from ao
float so = clamp(pow(NoV + ao, exp2(-16.0 * rough - 1.0)) - 1.0 + ao, 0.0, 1.0);
// check for NaN before min(): on this GPU min(NaN, x) returns x, which would hide the fault as a hot pixel
vec3 c = direct + (diffuseIBL * ao + specIBL * so) + bounce * ao + (fireLight(wpos, n, albedo) + handLight(wpos, n, albedo)) * ao;
if (any(isnan(direct)) || any(isnan(diffuseIBL)) || any(isnan(specIBL)) || isnan(so)) return vec3(0.0);
return sane(min(c, vec3(4096.0)));
}
// aerial perspective: exponential height fog toward the horizon sky, with sun inscatter
vec3 applyFog(vec3 col, vec3 wpos, float dist) {
vec3 dir = normalize(wpos - u_cam_pos);
float hf = u_fog_height_falloff;
float t = dir.y * hf;
float integ = (abs(t) > 1e-4) ? (1.0 - exp(-dist * t)) / t : dist * (1.0 - 0.5 * dist * t);
float fogAmt = u_fog_density * exp(-u_cam_pos.y * hf) * integ;
float f = 1.0 - exp(-fogAmt);
vec3 fogCol = skyPrefiltered(vec3(dir.x, max(dir.y, 0.02), dir.z), 0.6);
float sunAmt = pow(max(dot(dir, u_sun_dir), 0.0), 8.0);
fogCol += u_sun_color * 0.02 * sunAmt;
return mix(col, fogCol, clamp(f, 0.0, 1.0));
}

View file

@ -0,0 +1,185 @@
in vec3 v_wpos;
in vec3 v_nrm;
in vec2 v_uv;
in float v_seed;
in vec2 v_rot;
in float v_hull;
uniform float u_model_h;
out vec4 o_color;
uniform sampler2D u_diff;
uniform sampler2D u_nrm;
uniform sampler2D u_arm;
uniform mat4 u_view;
uniform vec3 u_tint;
uniform float u_rough_scale;
uniform float u_emissive; // self-lit (a flame): albedo added back after shading
#ifdef BLADE
uniform vec3 u_blade_base;
uniform vec3 u_blade_tip;
#endif
#ifdef CARD
uniform float u_cull; // the layer's cull distance (m); 0 = none
#endif
mat3 cotangentFrame(vec3 N, vec3 p, vec2 uv) {
vec3 dp1 = dFdx(p), dp2 = dFdy(p);
vec2 duv1 = dFdx(uv), duv2 = dFdy(uv);
vec3 dp2perp = cross(dp2, N), dp1perp = cross(N, dp1);
vec3 T = dp2perp * duv1.x + dp1perp * duv2.x;
vec3 B = dp2perp * duv1.y + dp1perp * duv2.y;
float invmax = inversesqrt(max(dot(T, T), dot(B, B)) + 1e-12);
return mat3(T * invmax, B * invmax, N);
}
void main() {
if (v_wpos.y < u_clip_y) discard;
vec3 N = normalize(v_nrm);
if (!gl_FrontFacing && v_hull >= 0.0) N = -N;
#ifdef CARD
// a baked card: albedo with coverage (premultiplied mips), normal in the card's own frame
vec4 ca = texture(u_diff, v_uv);
float rawA = ca.a;
ca.rgb /= max(ca.a, 1e-3);
#ifdef SHADOW_PASS
if (ca.a < 0.3) discard;
float cardAlpha = 1.0;
return;
#else
float cov = (ca.a - 0.4) / max(fwidth(ca.a), 1e-4) + 0.5;
if (cov < 0.02) discard;
float cardAlpha = clamp(cov, 0.0, 1.0);
#endif
#endif
float dist = length(v_wpos - u_cam_pos);
float viewDepth = -(u_view * vec4(v_wpos, 1.0)).z;
#if defined(CARD) && !defined(SHADOW_PASS)
// A clump card at 400 m is a two-pixel disc of pure leaf colour on whatever the ground
// is doing — on dark scree it glows. Dissolve toward the cull distance: the carpet
// drape on the ground carries the meadow's colour from there on.
if (u_cull > 1.0) cardAlpha *= 1.0 - smoothstep(u_cull * 0.55, u_cull, dist);
if (cardAlpha < 0.02) discard;
#endif
vec3 alb;
vec3 n;
vec3 arm;
float meshAlpha = 1.0;
#ifdef BLADE
// a procedural blade: dark at the root, lighter at the tip; some blades gone to seed
float t = clamp(v_uv.y, 0.0, 1.0);
alb = mix(u_blade_base, u_blade_tip, t * t) * regionTint(v_wpos, 0.35);
float dry = smoothstep(0.7, 0.8, fract(v_seed * 3.17));
alb = mix(alb, vec3(0.40, 0.36, 0.13) * (0.45 + 0.55 * t), dry * 0.75);
float patchy = fbm(v_wpos.xz * 0.045, 3) * 0.5 + 0.5;
alb *= mix(vec3(0.7, 0.8, 0.55), vec3(1.1, 1.05, 0.85), patchy);
// a far tuft is a patch of the meadow, darker than a lit blade tip and never straw
if (v_hull < 0.0) alb = mix(u_blade_base, u_blade_tip, 0.45) * regionTint(v_wpos, 0.35) * mix(vec3(0.7, 0.8, 0.55), vec3(1.1, 1.05, 0.85), patchy) * 0.72;
// a rounded cross-section reads softer than a flat card
vec3 side = normalize(cross(N, vec3(0.0, 1.0, 0.0)) + vec3(1e-4));
n = (v_hull < 0.0) ? N : normalize(N + side * (v_uv.x * 2.0 - 1.0) * 0.6);
arm = vec3(mix(0.2, 1.0, t * t), 0.85, 0.0);
#elif defined(CARD)
// thin grass is lit from either side: face the card toward the sun before shading
if (dot(N, u_sun_dir) < 0.0) N = -N;
// the normal atlas is premultiplied by coverage like the albedo (see impostor.frag)
vec4 cn = texture(u_nrm, v_uv) / max(rawA, 1e-3);
cn = clamp(cn, 0.0, 1.0);
vec3 bn = cn.rgb * 2.0 - 1.0;
// the card frame: right along the quad, up, and out of the quad
vec3 T = normalize(cross(vec3(0.0, 1.0, 0.0), N) + vec3(1e-5));
n = normalize(T * -bn.x + vec3(0.0, 1.0, 0.0) * bn.y + N * max(abs(bn.z), 0.25));
n = normalize(mix(n, normalize(N + vec3(0.0, 0.8, 0.0)), 0.35));
// A clump is a few pixels at 100 m: what the eye reads there is the hillside's shading,
// and a card facing the sun on a slope facing away from it glows against the ground
// like a sticker. Light it with the ground's own normal as it recedes (as the blades
// already do), so cover and terrain darken together.
#ifdef CHEAP
n = terrainNormalAt(v_wpos);
#else
n = normalize(mix(n, terrainNormalAt(v_wpos), smoothstep(25.0, 90.0, dist)));
#endif
alb = ca.rgb * 1.05 * regionTint(v_wpos, 0.8);
arm = vec3(mix(0.5, 1.0, clamp(v_uv.y, 0.0, 1.0)) * (0.6 + 0.4 * cn.a), 0.85, 0.0);
#elif defined(FLOWER)
// a lupine: green stem (uv.x < 1), violet florets above (uv.x in [1,2]), tinted per plant
float hue = fract(v_seed * 5.71);
float vy = clamp(v_uv.y, 0.0, 1.0);
vec3 violet = mix(vec3(0.07, 0.03, 0.32), vec3(0.28, 0.06, 0.36), hue);
vec3 tipc = mix(violet, vec3(0.5, 0.3, 0.7), 0.3);
if (v_uv.x >= 2.0) {
alb = vec3(0.05, 0.13, 0.025) * (0.7 + 0.6 * vy);
} else if (v_uv.x >= 1.0) {
float f = fract(v_uv.x);
alb = mix(violet, tipc, vy) * (0.65 + 0.35 * abs(f * 2.0 - 1.0));
// florets as little lobes: darker between them, a paler lip on each
float lobe = 0.55 + 0.45 * abs(sin(vy * 9.0 + f * 6.0));
alb = mix(alb * lobe, vec3(0.55, 0.45, 0.75), 0.18 * smoothstep(0.6, 1.0, lobe));
} else {
alb = vec3(0.07, 0.16, 0.03);
}
vec3 side = normalize(cross(N, vec3(0.0, 1.0, 0.0)) + vec3(1e-4));
n = normalize(N + side * (fract(v_uv.x) * 2.0 - 1.0) * 0.7);
arm = vec3(0.9, 0.7, 0.0);
#else
vec4 d = texture(u_diff, v_uv);
#ifdef ALPHA_TEST
// A needle sprig's alpha averages away in the mips, so a plain 0.5 test strips the
// crown bare past 50 m. Scale the alpha back up by the mip level (Castano's alpha
// mipmaps, done at sample time) and sharpen the edge with its screen derivative, then
// let alpha-to-coverage resolve it.
float lod = textureQueryLod(u_diff, v_uv).x;
float a = min(d.a * (1.0 + 0.45 * max(lod, 0.0)), 1.0);
float cov = (a - 0.4) / max(fwidth(a), 1e-4) + 0.5;
if (cov < 0.02) discard;
meshAlpha = clamp(cov, 0.0, 1.0);
#endif
vec3 tn = texture(u_nrm, v_uv).rgb * 2.0 - 1.0;
mat3 tbn = cotangentFrame(N, v_wpos, v_uv);
n = normalize(tbn * tn);
arm = texture(u_arm, v_uv).rgb;
n = normalize(mix(n, N, smoothstep(30.0, 120.0, dist)));
alb = d.rgb;
// a crown's interior is occluded by its own cards
if (u_model_h > 2.0) arm.r *= mix(0.5, 1.0, v_hull);
#endif
alb *= u_tint * (0.85 + 0.3 * fract(v_seed * 7.13));
#ifdef CARD
// a card is its own caster: look up the shadow a little above and in front of it, and let
// light bleed through the thin clump as real grass does
#ifdef CHEAP
float shadow = cloudShadow(v_wpos) * terrainShadow(v_wpos);
#else
float shadow = mix(1.0, sunShadow(v_wpos + N * 0.1 + vec3(0.0, 0.2, 0.0), N, viewDepth), 0.55);
#endif
#else
float shadow = sunShadow(v_wpos, N, viewDepth);
#endif
#ifdef FOLIAGE
// leaves and needles are matte at every angle: no grazing Fresnel on a two-sided card
float roughF = 1.0;
#else
float roughF = clamp(arm.g * u_rough_scale, 0.35, 1.0);
#endif
vec3 col = shade(v_wpos, n, alb, roughF, 0.0, arm.r, shadow, viewDepth);
#ifdef FOLIAGE
// thin-leaf translucency: light leaking through toward the viewer, and a wrapped diffuse
vec3 v = normalize(u_cam_pos - v_wpos);
float back = pow(max(dot(-v, u_sun_dir), 0.0), 3.0);
float wrap = max(dot(N, u_sun_dir) * 0.5 + 0.5, 0.0);
// Only a thin leaf a few metres away is translucent. A clump card at 200 m is a whole
// bush in two pixels, and giving it the leaf's glow toward the sun painted the
// backlit hillsides with lime discs. The term fades out with distance.
float thin = 1.0 - smoothstep(30.0, 140.0, dist);
// a dense crown of cards is not a thin leaf: much less light comes through it
if (u_model_h > 2.0) thin *= 0.3;
col += alb * u_sun_color * (0.14 * back + 0.05 * wrap) * thin * shadow * cloudShadow(v_wpos);
col += alb * skyIrradiance(vec3(0, 1, 0)) * 0.12 * arm.r;
#endif
col += alb * u_emissive;
if (any(isnan(col))) col = vec3(0.0);
col = applyFog(col, v_wpos, dist);
#ifdef CARD
o_color = vec4(sane(col), cardAlpha);
#else
o_color = vec4(sane(col), meshAlpha);
#endif
}

View file

@ -0,0 +1,102 @@
// instanced glTF model: attribute 3 = (x, y, z, scale), 4 = (sin yaw, cos yaw, seed, wind)
layout(location = 0) in vec3 a_pos;
layout(location = 1) in vec3 a_nrm;
layout(location = 2) in vec2 a_uv;
layout(location = 3) in vec4 i_pos;
layout(location = 4) in vec4 i_rot;
uniform mat4 u_view;
uniform mat4 u_proj;
uniform mat4 u_light_vp;
uniform float u_time;
uniform float u_wind;
uniform float u_card_w;
uniform float u_card_h;
#ifdef BLADE
uniform vec3 u_cam_pos;
uniform float u_cull; // the blade ring's edge (m)
#endif
// Ground cover is placed on the CPU from a bilinear read of the 4 m height texels, but
// the terrain is drawn from a B-spline of the same texels — two different surfaces,
// up to half a metre apart on rough ground, which buried blades and floated cards.
// Cover layers (u_ground) read the surface the terrain actually draws, so they always
// stand on it, at any tessellation level, with no hand-tuned lift.
uniform float u_ground;
uniform sampler2D u_ts_height;
uniform vec2 u_ts_origin;
uniform float u_ts_half;
float heightSmooth(sampler2D tex, vec2 uv) {
vec2 res = vec2(textureSize(tex, 0));
vec2 t = uv * res - 0.5;
vec2 f = fract(t);
vec2 i = floor(t);
vec2 w0 = (1.0 - f) * (1.0 - f) * (1.0 - f) / 6.0;
vec2 w1 = (4.0 - 6.0 * f * f + 3.0 * f * f * f) / 6.0;
vec2 w3 = f * f * f / 6.0;
vec2 w2 = 1.0 - w0 - w1 - w3;
vec2 s0 = w0 + w1, s1 = w2 + w3;
vec2 o0 = (i - 1.0 + w1 / s0 + 0.5) / res;
vec2 o1 = (i + 1.0 + w3 / s1 + 0.5) / res;
return (texture(tex, vec2(o0.x, o0.y)).r * s0.x + texture(tex, vec2(o1.x, o0.y)).r * s1.x) * s0.y
+ (texture(tex, vec2(o0.x, o1.y)).r * s0.x + texture(tex, vec2(o1.x, o1.y)).r * s1.x) * s1.y;
}
out vec3 v_wpos;
out vec3 v_nrm;
out vec2 v_uv;
out float v_seed;
out vec2 v_rot;
// Crown hull: a tree's needle cards are lit as if they were the surface of a rounded
// crown (normal from the crown's centre), not each as a flat top-lit quad, and the
// cards near the trunk are darkened as the crown's interior. This is how game trees
// have been shaded since SpeedTree; without it a card crown reads as frosted.
uniform float u_model_h; // the model's height (m), 0 = not a crown
out float v_hull; // 0 at the crown's axis .. 1 at its rim
void main() {
float s = i_rot.x, c = i_rot.y;
vec3 p = a_pos * i_pos.w;
#ifdef CARD
p = vec3(p.x * u_card_w, p.y * u_card_h, p.z * u_card_w);
#endif
vec3 n = a_nrm;
#ifdef BLADE
// distant blades: wider so a thinner field keeps its coverage, lit like the ground
// they stand on, and sunk into the carpet texture at the ring's edge instead of popping
float bd = distance(u_cam_pos.xz, i_pos.xz);
p.x *= 1.0 + 2.5 * smoothstep(12.0, 90.0, bd);
p.y *= 1.0 - smoothstep(u_cull * 0.72, u_cull, bd);
n = normalize(mix(n, vec3(0.0, 1.0, 0.0), smoothstep(15.0, 70.0, bd)));
#endif
v_rot = vec2(s, c);
p = vec3(c * p.x + s * p.z, p.y, -s * p.x + c * p.z);
n = vec3(c * n.x + s * n.z, n.y, -s * n.x + c * n.z);
#ifdef WIND
// sway grows with height above the base; gust phase from the instance seed
float hgt = max(p.y, 0.0);
float ph = u_time * 1.7 + i_rot.z * 6.2831 + i_pos.x * 0.05 + i_pos.z * 0.07;
float sway = (sin(ph) * 0.6 + sin(ph * 2.3 + 1.0) * 0.4) * u_wind * i_rot.w;
p.x += sway * hgt * hgt * 0.35;
p.z += sway * hgt * hgt * 0.15 * cos(ph * 0.7);
#endif
v_hull = 1.0;
if (u_model_h > 2.0) {
vec3 cc = vec3(0.0, u_model_h * i_pos.w * 0.55, 0.0);
vec3 rel = p - cc;
float rr = length(rel.xz) / max(u_model_h * i_pos.w * 0.28, 0.1);
v_hull = clamp(rr, 0.0, 1.0);
vec3 hull = normalize(vec3(rel.x, rel.y * 0.5, rel.z) + vec3(0.0, 0.15, 0.0));
n = normalize(mix(n, hull, 0.7));
}
vec3 w = p + i_pos.xyz;
if (u_ground > 0.5) {
vec2 huv = (i_pos.xz - u_ts_origin) / (2.0 * u_ts_half) + 0.5;
w.y = heightSmooth(u_ts_height, huv) - 0.03 + p.y;
}
v_wpos = w;
v_nrm = n;
v_uv = a_uv;
v_seed = i_rot.z;
#ifdef SHADOW_PASS
gl_Position = u_light_vp * vec4(w, 1.0);
#else
gl_Position = u_proj * u_view * vec4(w, 1.0);
#endif
}

View file

@ -0,0 +1,27 @@
// ---- shared noise (value / gradient / fbm / ridged) ----------------------------
float hash1(vec2 p) { return fract(sin(dot(p, vec2(127.1, 311.7))) * 43758.5453123); }
vec2 hash2(vec2 p) { p = vec2(dot(p, vec2(127.1, 311.7)), dot(p, vec2(269.5, 183.3))); return fract(sin(p) * 43758.5453123) * 2.0 - 1.0; }
float gnoise(vec2 p) {
vec2 i = floor(p), f = fract(p);
vec2 u = f * f * (3.0 - 2.0 * f);
return mix(mix(dot(hash2(i + vec2(0, 0)), f - vec2(0, 0)), dot(hash2(i + vec2(1, 0)), f - vec2(1, 0)), u.x),
mix(dot(hash2(i + vec2(0, 1)), f - vec2(0, 1)), dot(hash2(i + vec2(1, 1)), f - vec2(1, 1)), u.x), u.y);
}
float fbm(vec2 p, int oct) {
float a = 0.5, s = 0.0, n = 0.0;
mat2 r = mat2(0.8, 0.6, -0.6, 0.8) * 2.02;
for (int i = 0; i < oct; i++) { s += a * gnoise(p); n += a; a *= 0.5; p = r * p; }
return s / n;
}
float ridged(vec2 p, int oct) {
float a = 0.5, s = 0.0, w = 1.0;
mat2 r = mat2(0.8, 0.6, -0.6, 0.8) * 2.1;
for (int i = 0; i < oct; i++) { float n = 1.0 - abs(gnoise(p)); n = n * n * w; w = clamp(n * 1.5, 0.0, 1.0); s += a * n; a *= 0.5; p = r * p; }
return s;
}
// the dirt track: distance from a winding curve through the meadow
float pathDist(vec2 xz) {
float cx = 40.0 * sin(xz.y * 0.011) + 18.0 * sin(xz.y * 0.031 + 1.7) - 30.0;
float cx2 = -180.0 + 25.0 * sin(xz.y * 0.017 + 0.4) + (xz.y * 0.35);
return min(abs(xz.x - cx), abs(xz.x - cx2) + 1.0);
}

View file

@ -0,0 +1,14 @@
// 2D overlay: a texture times a colour; the font atlas is white glyphs on alpha
in vec2 v_uv;
in vec4 v_col;
uniform sampler2D u_tex;
uniform float u_is_font;
out vec4 o_color;
void main() {
vec4 t = texture(u_tex, v_uv);
if (u_is_font > 0.5) {
o_color = vec4(v_col.rgb, v_col.a * t.a);
} else {
o_color = vec4(v_col.rgb * t.rgb, v_col.a * t.a);
}
}

View file

@ -0,0 +1,13 @@
// 2D overlay (overlay.ludic): pixels with the origin top-left, straight to clip space
layout(location = 0) in vec2 a_pos;
layout(location = 1) in vec2 a_uv;
layout(location = 2) in vec4 a_col;
uniform vec2 u_screen;
out vec2 v_uv;
out vec4 v_col;
void main() {
vec2 p = a_pos / u_screen * 2.0 - 1.0;
gl_Position = vec4(p.x, -p.y, 0.0, 1.0);
v_uv = a_uv;
v_col = a_col;
}

View file

@ -0,0 +1,11 @@
#ifdef ALPHA_TEST
// foliage meshes are cut-out cards: their shadow must have the card's shape, not the quad's
in vec2 v_uv;
uniform sampler2D u_diff;
void main() {
float lod = textureQueryLod(u_diff, v_uv).x;
if (texture(u_diff, v_uv).a * (1.0 + 0.45 * max(lod, 0.0)) < 0.45) discard;
}
#else
void main() { }
#endif

View file

@ -0,0 +1,21 @@
// luma unsharp mask + film grain, on the final LDR image
in vec2 v_uv;
out vec4 o_color;
uniform sampler2D u_src;
uniform vec2 u_texel;
uniform float u_amount;
uniform float u_grain;
float hash(vec2 p) { return fract(sin(dot(p, vec2(12.9898, 78.233))) * 43758.5453); }
float luma(vec3 c) { return dot(c, vec3(0.299, 0.587, 0.114)); }
void main() {
vec3 c = texture(u_src, v_uv).rgb;
vec3 n = texture(u_src, v_uv + vec2(0, u_texel.y)).rgb, s = texture(u_src, v_uv - vec2(0, u_texel.y)).rgb;
vec3 e = texture(u_src, v_uv + vec2(u_texel.x, 0)).rgb, w = texture(u_src, v_uv - vec2(u_texel.x, 0)).rgb;
float lc = luma(c);
float lb = (luma(n) + luma(s) + luma(e) + luma(w) + lc * 4.0) / 8.0;
float d = clamp((lc - lb) * u_amount, -0.08, 0.08);
vec3 col = c * (1.0 + d / max(lc, 1e-3));
float g = (hash(v_uv * 1731.0 + fract(u_time)) - 0.5) * u_grain;
col += g * (0.6 + 0.4 * (1.0 - lc));
o_color = vec4(clamp(col, 0.0, 1.0), 1.0);
}

View file

@ -0,0 +1,43 @@
// a skinned glTF model (skin.ludic / actor.ludic): four joint influences per vertex
// blended on the GPU, then one model matrix. Writes the same varyings as model.vert so
// model.frag (lit) and shadow.frag (casters) shade it unchanged.
layout(location = 0) in vec3 a_pos;
layout(location = 1) in vec3 a_nrm;
layout(location = 2) in vec2 a_uv;
layout(location = 5) in vec4 a_joints; // integer indices, read as floats
layout(location = 6) in vec4 a_weights;
uniform mat4 u_model;
uniform mat4 u_bones[48];
uniform float u_skinned; // 0: a rigid model on the same path
uniform mat4 u_view;
uniform mat4 u_proj;
uniform mat4 u_light_vp;
out vec3 v_wpos;
out vec3 v_nrm;
out vec2 v_uv;
out float v_seed;
out vec2 v_rot;
out float v_hull;
void main() {
mat4 m = u_model;
if (u_skinned > 0.5) {
mat4 sk = a_weights.x * u_bones[int(a_joints.x + 0.5)]
+ a_weights.y * u_bones[int(a_joints.y + 0.5)]
+ a_weights.z * u_bones[int(a_joints.z + 0.5)]
+ a_weights.w * u_bones[int(a_joints.w + 0.5)];
m = u_model * sk;
}
vec4 w = m * vec4(a_pos, 1.0);
v_wpos = w.xyz;
v_nrm = normalize(mat3(m) * a_nrm);
v_uv = a_uv;
// model.frag scales the albedo by 0.85 + 0.3 * fract(seed * 7.13); this seed makes that 1
v_seed = 0.0701;
v_rot = vec2(0.0, 1.0);
v_hull = 1.0;
#ifdef SHADOW_PASS
gl_Position = u_light_vp * w;
#else
gl_Position = u_proj * u_view * w;
#endif
}

View file

@ -0,0 +1,40 @@
in vec2 v_uv;
out vec4 o_color;
uniform sampler2D u_sky;
uniform mat4 u_inv_vp;
uniform float u_sky_gain;
uniform float u_sky_sat;
// a star: one hash per cell of the direction, a few of them bright, a slow twinkle
float starField(vec3 dir, float t) {
vec3 p = dir * 230.0;
vec3 c = floor(p);
vec3 f = p - c - 0.5;
float h = fract(sin(dot(c, vec3(12.9898, 78.233, 37.719))) * 43758.5453);
float h2 = fract(h * 91.7);
float bright = smoothstep(0.972, 1.0, h);
float disc = smoothstep(0.42, 0.0, length(f));
float twinkle = 0.7 + 0.3 * sin(t * (1.5 + 3.0 * h2) + h2 * 40.0);
return bright * disc * twinkle * (0.5 + h2);
}
void main() {
vec4 a = u_inv_vp * vec4(v_uv * 2.0 - 1.0, 1.0, 1.0);
vec3 dir = normalize(a.xyz / a.w - u_cam_pos);
// level 0: the equirect seam (atan wraps) would otherwise pick the smallest mip along one column
vec3 col = min(textureLod(u_sky, skyUV(dir), 0.0).rgb, vec3(4096.0)) * u_sky_gain;
float l = dot(col, vec3(0.2126, 0.7152, 0.0722));
col = max(mix(vec3(l), col, u_sky_sat), vec3(0.0));
// the photograph's sky dims with the day (daylight.ludic); the night adds its own
col *= u_ibl_scale;
float night = 1.0 - smoothstep(0.0, 0.45, u_daylight);
if (night > 0.0) {
vec3 nightCol = mix(vec3(0.012, 0.016, 0.034), vec3(0.003, 0.004, 0.010), clamp(dir.y, 0.0, 1.0));
float stars = starField(dir, u_time) * smoothstep(-0.02, 0.15, dir.y);
nightCol += stars * vec3(0.55, 0.6, 0.7) * night;
col += nightCol * night;
}
// below the horizon the HDRI ground is replaced by the fog colour
float below = smoothstep(0.0, -0.08, dir.y);
vec3 fogCol = skyPrefiltered(vec3(dir.x, 0.02, dir.z), 0.6);
col = mix(col, fogCol, below);
o_color = vec4(sane(col), 1.0);
}

View file

@ -0,0 +1,47 @@
// screen-space ambient occlusion from the resolved depth (half resolution)
in vec2 v_uv;
out vec4 o_color;
uniform sampler2D u_depth;
uniform mat4 u_inv_proj;
uniform mat4 u_proj;
uniform vec2 u_texel;
uniform float u_radius; // world metres
uniform float u_intensity;
vec3 viewPos(vec2 uv) {
float d = texture(u_depth, uv).r;
vec4 p = u_inv_proj * vec4(uv * 2.0 - 1.0, d * 2.0 - 1.0, 1.0);
return p.xyz / p.w;
}
void main() {
vec3 P = viewPos(v_uv);
if (-P.z > 900.0) { o_color = vec4(1.0); return; }
// normal from the depth's neighbourhood (take the smaller difference on each axis)
vec3 Pr = viewPos(v_uv + vec2(u_texel.x, 0.0)), Pl = viewPos(v_uv - vec2(u_texel.x, 0.0));
vec3 Pu = viewPos(v_uv + vec2(0.0, u_texel.y)), Pd = viewPos(v_uv - vec2(0.0, u_texel.y));
vec3 dx = (abs(Pr.z - P.z) < abs(P.z - Pl.z)) ? Pr - P : P - Pl;
vec3 dy = (abs(Pu.z - P.z) < abs(P.z - Pd.z)) ? Pu - P : P - Pd;
vec3 N = normalize(cross(dx, dy));
float noise = ign(gl_FragCoord.xy);
float ao = 0.0;
const int S = 12;
float radius = u_radius * (1.0 + 0.01 * -P.z);
for (int i = 0; i < S; i++) {
float a = (float(i) + noise) * 2.3999632; // golden angle spiral
float r = sqrt((float(i) + 0.5 + noise) / float(S));
vec3 dir = vec3(cos(a) * r, sin(a) * r, sqrt(max(0.0, 1.0 - r * r)));
// hemisphere around N
vec3 up = abs(N.z) < 0.999 ? vec3(0, 0, 1) : vec3(1, 0, 0);
vec3 t = normalize(cross(up, N)), b = cross(N, t);
vec3 s = P + (t * dir.x + b * dir.y + N * dir.z) * radius * (0.2 + 0.8 * r);
vec4 c = u_proj * vec4(s, 1.0);
vec2 suv = c.xy / c.w * 0.5 + 0.5;
if (suv.x < 0.0 || suv.x > 1.0 || suv.y < 0.0 || suv.y > 1.0) continue;
float sz = viewPos(suv).z;
float rangeCheck = smoothstep(0.0, 1.0, radius / max(abs(P.z - sz), 1e-3));
ao += (sz >= s.z + 0.02 * radius ? 1.0 : 0.0) * rangeCheck;
}
ao = 1.0 - u_intensity * ao / float(S);
// contact occlusion is a near-field effect: fade it out with distance
ao = mix(clamp(ao, 0.0, 1.0), 1.0, smoothstep(120.0, 350.0, -P.z));
o_color = vec4(ao, -P.z, 0.0, 1.0);
}

View file

@ -0,0 +1,19 @@
// depth-aware 4x4 blur of the ao (a) and indirect bounce (rgb)
in vec2 v_uv;
out vec4 o_color;
uniform sampler2D u_ao;
uniform sampler2D u_depth;
uniform vec2 u_texel;
void main() {
float cd = texture(u_depth, v_uv).r;
vec4 sum = vec4(0.0);
float wsum = 0.0;
for (int y = -2; y < 2; y++) for (int x = -2; x < 2; x++) {
vec2 o = vec2(float(x) + 0.5, float(y) + 0.5) * u_texel;
vec4 s = texture(u_ao, v_uv + o);
float sd = texture(u_depth, v_uv + o).r;
float w = exp(-abs(sd - cd) * 4000.0);
sum += s * w; wsum += w;
}
o_color = sum / max(wsum, 1e-4);
}

View file

@ -0,0 +1,66 @@
// screen-space ambient occlusion + one indirect diffuse bounce (SSGI) from the previous
// frame's lit colour; half resolution, denoised over time by the TAA history it feeds
in vec2 v_uv;
out vec4 o_color;
uniform sampler2D u_depth;
uniform sampler2D u_prev_color; // last frame's anti-aliased HDR colour
uniform mat4 u_inv_proj;
uniform mat4 u_proj;
uniform vec2 u_texel;
uniform float u_radius;
uniform float u_intensity;
uniform float u_frame;
vec3 viewPos(vec2 uv) {
float d = texture(u_depth, uv).r;
vec4 p = u_inv_proj * vec4(uv * 2.0 - 1.0, d * 2.0 - 1.0, 1.0);
return p.xyz / p.w;
}
void main() {
vec3 P = viewPos(v_uv);
if (-P.z > 900.0) { o_color = vec4(0.0, 0.0, 0.0, 1.0); return; }
vec3 Pr = viewPos(v_uv + vec2(u_texel.x, 0.0)), Pl = viewPos(v_uv - vec2(u_texel.x, 0.0));
vec3 Pu = viewPos(v_uv + vec2(0.0, u_texel.y)), Pd = viewPos(v_uv - vec2(0.0, u_texel.y));
vec3 dx = (abs(Pr.z - P.z) < abs(P.z - Pl.z)) ? Pr - P : P - Pl;
vec3 dy = (abs(Pu.z - P.z) < abs(P.z - Pd.z)) ? Pu - P : P - Pd;
vec3 N = normalize(cross(dx, dy));
// A fixed per-pixel dither, not a per-frame one. Advancing the sequence every frame
// spreads the sampling error over time, which is only an improvement if something
// then averages the frames; with no temporal anti-aliasing left it is just noise that
// changes every frame, and it was the largest single source of the flicker on movement.
float noise = ign(gl_FragCoord.xy);
float ao = 0.0;
vec3 gi = vec3(0.0);
float giW = 0.0;
const int S = 8;
float radius = u_radius * (1.0 + 0.01 * -P.z);
vec3 up = abs(N.z) < 0.999 ? vec3(0, 0, 1) : vec3(1, 0, 0);
vec3 t = normalize(cross(up, N)), b = cross(N, t);
for (int i = 0; i < S; i++) {
float a = (float(i) + noise) * 2.3999632;
float r = sqrt((float(i) + 0.5 + noise) / float(S));
vec3 dir = vec3(cos(a) * r, sin(a) * r, sqrt(max(0.0, 1.0 - r * r)));
vec3 wdir = t * dir.x + b * dir.y + N * dir.z;
vec3 s = P + wdir * radius * (0.2 + 0.8 * r);
vec4 c = u_proj * vec4(s, 1.0);
vec2 suv = c.xy / c.w * 0.5 + 0.5;
if (suv.x < 0.0 || suv.x > 1.0 || suv.y < 0.0 || suv.y > 1.0) continue;
vec3 sp = viewPos(suv);
float rangeCheck = smoothstep(0.0, 1.0, radius / max(abs(P.z - sp.z), 1e-3));
bool occluded = sp.z >= s.z + 0.02 * radius;
ao += (occluded ? 1.0 : 0.0) * rangeCheck;
// the occluder's lit colour bounces back toward P (weighted by how squarely it faces P)
if (occluded) {
vec3 toS = sp - P;
float d2 = max(dot(toS, toS), 1e-3);
float cosP = max(dot(N, toS) * inversesqrt(d2), 0.0);
vec3 col = sane(texture(u_prev_color, suv).rgb);
gi += col * cosP * rangeCheck;
giW += 1.0;
}
}
ao = 1.0 - u_intensity * ao / float(S);
float fade = smoothstep(120.0, 350.0, -P.z);
ao = mix(clamp(ao, 0.0, 1.0), 1.0, fade);
gi = (giW > 0.0 ? gi / float(S) : vec3(0.0)) * (1.0 - fade);
o_color = vec4(gi, ao);
}

View file

@ -0,0 +1,32 @@
// Second generation pass: copy the height into R and bake the B-spline surface normal into
// GBA, once, at texel resolution. terrain.frag used to differentiate the bicubic height
// per pixel — four bicubic reads, sixteen taps — for a quantity that never changes.
in vec2 v_uv;
out vec4 o;
uniform sampler2D u_src;
uniform float u_half;
float heightSmooth(sampler2D tex, vec2 uv) {
vec2 res = vec2(textureSize(tex, 0));
vec2 t = uv * res - 0.5;
vec2 f = fract(t);
vec2 i = floor(t);
vec2 w0 = (1.0 - f) * (1.0 - f) * (1.0 - f) / 6.0;
vec2 w1 = (4.0 - 6.0 * f * f + 3.0 * f * f * f) / 6.0;
vec2 w3 = f * f * f / 6.0;
vec2 w2 = 1.0 - w0 - w1 - w3;
vec2 s0 = w0 + w1, s1 = w2 + w3;
vec2 o0 = (i - 1.0 + w1 / s0 + 0.5) / res;
vec2 o1 = (i + 1.0 + w3 / s1 + 0.5) / res;
return (texture(tex, vec2(o0.x, o0.y)).r * s0.x + texture(tex, vec2(o1.x, o0.y)).r * s1.x) * s0.y
+ (texture(tex, vec2(o0.x, o1.y)).r * s0.x + texture(tex, vec2(o1.x, o1.y)).r * s1.x) * s1.y;
}
void main() {
float step = 1.0 / float(textureSize(u_src, 0).x);
float world = step * 2.0 * u_half;
float hl = heightSmooth(u_src, v_uv - vec2(step, 0));
float hr = heightSmooth(u_src, v_uv + vec2(step, 0));
float hd = heightSmooth(u_src, v_uv - vec2(0, step));
float hu = heightSmooth(u_src, v_uv + vec2(0, step));
vec3 n = normalize(vec3(hl - hr, 2.0 * world, hd - hu));
o = vec4(texture(u_src, v_uv).r, n);
}

View file

@ -0,0 +1,467 @@
in vec3 v_wpos;
in vec2 v_huv;
out vec4 o_color;
#ifdef TFAST_2
#define fbm(p, o) 0.1
#define ridged(p, o) 0.3
#define gnoise(p) 0.1
#endif
uniform sampler2D u_height;
uniform float u_half;
uniform float u_texel; // height-map texel size in uv
uniform sampler2D u_grass_d; uniform sampler2D u_grass_n; uniform sampler2D u_grass_a;
uniform sampler2D u_ortho;
uniform sampler2D u_sunshadow; // the sun visibility this pixel already has (tersun.frag)
uniform float u_ortho_on;
uniform sampler2D u_rock_d; uniform sampler2D u_rock_n; uniform sampler2D u_rock_a;
uniform sampler2D u_snow_d; uniform sampler2D u_carpet; // the clump cards baked straight down (alpha = coverage)
uniform float u_carpet_on;
uniform float u_snow_line;
uniform float u_lake_level; // the ground just above the water is wet and dark
uniform vec2 u_origin; // world offset of the terrain grid
// The ground's own sun shadow comes from the baked height-field map (tershadow.frag),
// applied inside sunShadow() for every receiver in the scene.
// bicubic (B-spline) sample through four bilinear taps: the 10 m photo pixels stop reading as squares
vec3 orthoSmooth(vec2 uv) {
vec2 res = vec2(textureSize(u_ortho, 0));
vec2 t = uv * res - 0.5;
vec2 f = fract(t);
vec2 i = floor(t);
vec2 w0 = (1.0 - f) * (1.0 - f) * (1.0 - f) / 6.0, w1 = (4.0 - 6.0 * f * f + 3.0 * f * f * f) / 6.0;
vec2 w3 = f * f * f / 6.0, w2 = 1.0 - w0 - w1 - w3;
vec2 s0 = w0 + w1, s1 = w2 + w3;
vec2 o0 = (i - 1.0 + w1 / s0 + 0.5) / res, o1 = (i + 1.0 + w3 / s1 + 0.5) / res;
return (texture(u_ortho, vec2(o0.x, o0.y)).rgb * s0.x + texture(u_ortho, vec2(o1.x, o0.y)).rgb * s1.x) * s0.y
+ (texture(u_ortho, vec2(o0.x, o1.y)).rgb * s0.x + texture(u_ortho, vec2(o1.x, o1.y)).rgb * s1.x) * s1.y;
}
uniform mat4 u_view;
// B-spline bicubic sample of the height field, through four bilinear taps. The height
// texture is only C0 under bilinear filtering: its slope jumps at every texel edge, and
// the mesh chords across each triangle, so the geometry and the normal were reading two
// different surfaces and the shading kinked along every triangle diagonal. Both stages
// call this, so they now agree on one smooth surface.
float heightSmooth(sampler2D tex, vec2 uv) {
vec2 res = vec2(textureSize(tex, 0));
vec2 t = uv * res - 0.5;
vec2 f = fract(t);
vec2 i = floor(t);
vec2 w0 = (1.0 - f) * (1.0 - f) * (1.0 - f) / 6.0;
vec2 w1 = (4.0 - 6.0 * f * f + 3.0 * f * f * f) / 6.0;
vec2 w3 = f * f * f / 6.0;
vec2 w2 = 1.0 - w0 - w1 - w3;
vec2 s0 = w0 + w1, s1 = w2 + w3;
vec2 o0 = (i - 1.0 + w1 / s0 + 0.5) / res;
vec2 o1 = (i + 1.0 + w3 / s1 + 0.5) / res;
return (texture(tex, vec2(o0.x, o0.y)).r * s0.x + texture(tex, vec2(o1.x, o0.y)).r * s1.x) * s0.y
+ (texture(tex, vec2(o0.x, o1.y)).r * s0.x + texture(tex, vec2(o1.x, o1.y)).r * s1.x) * s1.y;
}
// baked at generation (ternormal.frag) into the height texture's GBA
vec3 terrainNormal(vec2 uv) {
return normalize(texture(u_height, uv).gba);
}
// stochastic (triangle-grid) sampling: three randomly offset / rotated taps blended by
// barycentric weights, so a scanned tile never repeats visibly
void triGrid(vec2 uv, out float w1, out float w2, out float w3, out vec2 v1, out vec2 v2, out vec2 v3) {
const mat2 skew = mat2(1.0, 0.0, -0.57735027, 1.15470054);
vec2 sk = skew * (uv * 3.4641016);
vec2 base = floor(sk);
vec3 t = vec3(fract(sk), 0.0);
t.z = 1.0 - t.x - t.y;
if (t.z > 0.0) { w1 = t.z; w2 = t.y; w3 = t.x; v1 = base; v2 = base + vec2(0, 1); v3 = base + vec2(1, 0); }
else { w1 = -t.z; w2 = 1.0 - t.y; w3 = 1.0 - t.x; v1 = base + vec2(1, 1); v2 = base + vec2(1, 0); v3 = base + vec2(0, 1); }
}
// The per-cell rotation must be applied to the DERIVATIVES as well as the coordinate.
// Handing textureGrad the gradients of the unrotated uv makes every cell sample with a
// footprint pointing the wrong way, so each one lands on a slightly different mip and
// anisotropy — and that per-cell difference is exactly the faint lattice over every
// surface. Returning the rotation lets the caller transform its gradients to match.
mat2 cellRot(vec2 cell) {
float a = hash1(cell) * 6.2831853;
float c = cos(a), s = sin(a);
return mat2(c, s, -s, c);
}
vec2 rotUV(vec2 uv, vec2 cell) {
return cellRot(cell) * uv + hash2(cell + 3.7) * 4.0;
}
// A tap whose weight rounds away is a tap not worth taking. The barycentric weights are
// raised to the fourth power to sharpen the blend, which leaves one of the three
// dominant over most of the plane and the other two often at a few thousandths; taking
// only the ones that carry any of the result, and renormalising over those, is
// indistinguishable from taking all three and is most of what this shader used to spend
// on the ground. TRI_EPS is the weight below which a tap cannot move an 8-bit channel.
#define TRI_EPS 0.004
vec4 sampleCarpet(vec2 uv, vec2 dx, vec2 dy) {
float w1, w2, w3; vec2 v1, v2, v3;
triGrid(uv * 0.3, w1, w2, w3, v1, v2, v3);
vec3 w = pow(vec3(w1, w2, w3), vec3(4.0)); w /= (w.x + w.y + w.z);
vec4 acc = vec4(0.0);
float wsum = 0.0;
if (w.x > TRI_EPS) { mat2 C = cellRot(v1); acc += textureGrad(u_carpet, rotUV(uv, v1), C * dx, C * dy) * w.x; wsum += w.x; }
if (w.y > TRI_EPS) { mat2 C = cellRot(v2); acc += textureGrad(u_carpet, rotUV(uv, v2), C * dx, C * dy) * w.y; wsum += w.y; }
if (w.z > TRI_EPS) { mat2 C = cellRot(v3); acc += textureGrad(u_carpet, rotUV(uv, v3), C * dx, C * dy) * w.z; wsum += w.z; }
return acc / max(wsum, 1e-4);
}
// Stochastic (triangle-grid) sampling: three randomly offset / rotated taps blended by
// barycentric weights, so a scanned tile never repeats visibly. The gradients are
// rotated per cell to match each tap's own rotation — handing textureGrad the gradients
// of the unrotated uv makes every cell sample with a footprint pointing the wrong way,
// landing on a different mip and anisotropy.
// one cell of the triangle grid: its rotation, its offset, and its three maps
void matTap(sampler2D d, sampler2D nm, sampler2D am, vec2 uv, vec2 dx, vec2 dy, vec2 cell, float wt,
inout vec3 alb, inout vec3 nsum, inout vec3 arm, inout float wsum) {
mat2 R = cellRot(cell);
vec2 u = R * uv + hash2(cell + 3.7) * 4.0;
vec2 gx = R * dx, gy = R * dy;
alb += textureGrad(d, u, gx, gy).rgb * wt;
nsum += (textureGrad(nm, u, gx, gy).rgb * 2.0 - 1.0) * wt;
arm += textureGrad(am, u, gx, gy).rgb * wt;
wsum += wt;
}
// Stochastic (triangle-grid) sampling: three randomly offset / rotated taps blended by
// barycentric weights, so a scanned tile never repeats visibly. The gradients are
// rotated per cell to match each tap's own rotation — handing textureGrad the gradients
// of the unrotated uv makes every cell sample with a footprint pointing the wrong way,
// landing on a different mip and anisotropy.
//
// The weights are sharpened to the fourth power, which leaves one cell dominant over
// most of the plane and the other two at a few thousandths. Everything a cell needs —
// its rotation (a hash, a sine and a cosine), its offset, its two rotated gradients —
// is computed inside its own test, so a cell that cannot move the result costs nothing.
void sampleMat(sampler2D d, sampler2D nm, sampler2D am, vec2 uv, vec2 dx, vec2 dy, out vec3 alb, out vec3 nrm, out vec3 arm) {
float w1, w2, w3; vec2 v1, v2, v3;
triGrid(uv * 0.3, w1, w2, w3, v1, v2, v3);
vec3 w = pow(vec3(w1, w2, w3), vec3(4.0)); w /= (w.x + w.y + w.z);
alb = vec3(0.0); arm = vec3(0.0);
vec3 nsum = vec3(0.0);
float wsum = 0.0;
if (w.x > TRI_EPS) { matTap(d, nm, am, uv, dx, dy, v1, w.x, alb, nsum, arm, wsum); }
if (w.y > TRI_EPS) { matTap(d, nm, am, uv, dx, dy, v2, w.y, alb, nsum, arm, wsum); }
if (w.z > TRI_EPS) { matTap(d, nm, am, uv, dx, dy, v3, w.z, alb, nsum, arm, wsum); }
float iw = 1.0 / max(wsum, 1e-4);
alb *= iw; arm *= iw;
// rotate the tangent normals back with their taps
nrm = normalize(nsum);
}
void samplePlain(sampler2D d, sampler2D nm, sampler2D am, vec2 uv, vec2 dx, vec2 dy, out vec3 alb, out vec3 nrm, out vec3 arm) {
alb = textureGrad(d, uv, dx, dy).rgb;
nrm = textureGrad(nm, uv, dx, dy).rgb * 2.0 - 1.0;
arm = textureGrad(am, uv, dx, dy).rgb;
}
// triplanar sample for steep rock
void sampleTri(sampler2D d, sampler2D nm, sampler2D am, vec3 p, vec3 dpx, vec3 dpy, vec3 n, float scale, out vec3 alb, out vec3 nrm, out vec3 arm) {
vec3 w = pow(abs(n), vec3(4.0)); w /= (w.x + w.y + w.z);
// The fourth power leaves ground facing one axis almost entirely on that axis's plane:
// a slope has to be within a few degrees of a diagonal before a second projection
// carries anything, and the third almost never does.
vec3 a0, n0, r0;
alb = vec3(0.0); arm = vec3(0.0);
vec3 nsum = vec3(0.0);
float wsum = 0.0;
if (w.x > TRI_EPS) {
samplePlain(d, nm, am, p.zy * scale, dpx.zy * scale, dpy.zy * scale, a0, n0, r0);
alb += a0 * w.x; arm += r0 * w.x;
nsum += vec3(n0.xy + n.zy, abs(n0.z) * n.x).zyx * w.x;
wsum += w.x;
}
if (w.y > TRI_EPS) {
samplePlain(d, nm, am, p.xz * scale, dpx.xz * scale, dpy.xz * scale, a0, n0, r0);
alb += a0 * w.y; arm += r0 * w.y;
nsum += vec3(n0.xy + n.xz, abs(n0.z) * n.y).xzy * w.y;
wsum += w.y;
}
if (w.z > TRI_EPS) {
samplePlain(d, nm, am, p.xy * scale, dpx.xy * scale, dpy.xy * scale, a0, n0, r0);
alb += a0 * w.z; arm += r0 * w.z;
nsum += vec3(n0.xy + n.xy, abs(n0.z) * n.z) * w.z;
wsum += w.z;
}
float iw = 1.0 / max(wsum, 1e-4);
alb *= iw; arm *= iw;
nrm = normalize(nsum);
}
vec3 dbg_n; vec3 dbg_alb; float dbg_shadow; vec3 dbg_mat;
// cheap = the far tier: single taps, noise at its mean, one shadow tap. Same code, same
// mean colour, so the tier boundary cannot show as a ring.
float fbmC(bool cheap, vec2 q, int o) { return cheap ? 0.0 : fbm(q, o); }
float ridgedC(bool cheap, vec2 q, int o) { return cheap ? 0.35 : ridged(q, o); }
float gnoiseC(bool cheap, vec2 q) { return cheap ? 0.0 : gnoise(q); }
vec3 orthoC(bool cheap, vec2 uv) { return cheap ? textureLod(u_ortho, uv, 1.0).rgb : orthoSmooth(uv); }
vec4 carpetC(bool cheap, vec2 uv, vec2 dx, vec2 dy) { return cheap ? textureGrad(u_carpet, uv, dx, dy) : sampleCarpet(uv, dx, dy); }
void matC(bool cheap, sampler2D d, sampler2D nm, sampler2D am, vec2 uv, vec2 dx, vec2 dy, out vec3 alb, out vec3 nrm, out vec3 arm) {
if (cheap) samplePlain(d, nm, am, uv, dx, dy, alb, nrm, arm); else sampleMat(d, nm, am, uv, dx, dy, alb, nrm, arm);
}
void triC(bool cheap, sampler2D d, sampler2D nm, sampler2D am, vec3 p, vec3 dpx, vec3 dpy, vec3 n, float scale, out vec3 alb, out vec3 nrm, out vec3 arm) {
if (cheap) { samplePlain(d, nm, am, p.xz * scale, dpx.xz * scale, dpy.xz * scale, alb, nrm, arm); nrm = normalize(vec3(nrm.x, 1.0, nrm.y) + vec3(0.0, 1e-3, 0.0)); }
else sampleTri(d, nm, am, p, dpx, dpy, n, scale, alb, nrm, arm);
}
vec3 groundShade(vec3 p, vec3 N, float slope, float dist, float viewDepth, bool cheap) {
// ---- material weights ----
float macro = fbmC(cheap, p.xz * 0.02, 2);
// a foot track: two metres wide, worn into whatever the ground is, not a painted band
float pathW = 0.6 * smoothstep(3.0 + 0.8 * macro, 1.0, pathDist(p.xz)) * smoothstep(0.35, 0.1, slope);
float rockW = max(smoothstep(0.30, 0.55, slope + 0.1 * macro), 0.9 * smoothstep(170.0, 300.0, p.y + 30.0 * macro));
// The ridge field places the snow line's raggedness and nothing else. Both terms below
// are zero more than 160 m under the snow line whatever it returns (macro and ridgeN
// can lift the test height by at most 100 m), which is the whole valley floor.
float snowW = 0.0;
if (p.y > u_snow_line - 160.0) {
float ridgeN = ridgedC(cheap, p.xz * 0.0018 + 11.0, 2);
snowW = smoothstep(u_snow_line - 60.0, u_snow_line + 60.0, p.y + 60.0 * macro + 40.0 * ridgeN) * smoothstep(0.55, 0.15, slope);
// wind-packed snow lingers in the gullies of the steep faces too
snowW = max(snowW, 0.6 * smoothstep(u_snow_line - 120.0, u_snow_line, p.y) * smoothstep(0.45, 0.2, slope) * smoothstep(0.55, 0.75, ridgeN));
}
float grassW = 1.0 - max(pathW, max(rockW, snowW));
// ---- what the photograph says is here -------------------------------------------
// This classification used to sit between the material samples, which meant every
// pixel sampled every material before anything knew which of them it would use. It
// runs first now: it costs three filtered taps of the survey image and it decides
// whether the scanned grass, rock and snow are needed at all.
vec3 oc = vec3(0.0);
float forestW = 0.0; // dense conifer: the ground under it is duff, not meadow
float screeC = 0.0; // bare ground: talus, moraine gravel, the lake's cobble shore
float snowC = 0.0;
bool orthoOn = u_ortho_on > 0.5;
#ifdef TFAST_4
orthoOn = false;
#endif
if (orthoOn) {
oc = orthoC(cheap, v_huv);
// (classified from a ~40 m blur: thresholding the raw 10 m pixels drew hard squares)
vec3 ocf = textureLod(u_ortho, v_huv, 2.0).rgb;
float gx = ocf.g - max(ocf.r, ocf.b);
// (the photograph is sampled linear: sRGB 72 is 0.06, 92 is 0.11)
forestW = smoothstep(0.12, 0.06, max(ocf.r, max(ocf.g, ocf.b))) * smoothstep(0.004, 0.012, gx) * smoothstep(u_lake_level + 0.8, u_lake_level + 1.8, p.y);
// Classify from a ~60 m blur, never from the pixels: the survey's 10 m pixels carry
// a foot trail as a broken line of bare ground, and thresholding them painted it
// across the meadow as tan dashes (and, on the CPU, lined boulders up along it).
vec3 ocl = textureLod(u_ortho, v_huv, 2.5).rgb;
float mxc = max(ocl.r, max(ocl.g, ocl.b)), mnc = min(ocl.r, min(ocl.g, ocl.b));
float greenEx = ocl.g - max(ocl.r, ocl.b);
screeC = smoothstep(0.008, -0.002, greenEx) * smoothstep(0.06, 0.12, mxc) * (1.0 - smoothstep(0.55, 0.75, mxc)) * smoothstep(u_lake_level + 0.2, u_lake_level + 1.2, p.y);
snowC = smoothstep(0.08, 0.04, mxc - mnc) * smoothstep(0.55, 0.8, mxc);
}
// snow lingering in the high gullies is drawn further down, but whether it can be
// there at all is known now, and it is the third caller of the snow sample
float gullyGate = smoothstep(450.0, 650.0, p.y) * smoothstep(0.75, 0.35, slope);
// ---- samples ---------------------------------------------------------------------
// World-space derivatives, taken once and in unbranched control flow: every sample
// below is in a branch and takes its gradients from these.
vec3 dpx = dFdx(p), dpy = dFdy(p);
vec3 gA = vec3(0.3, 0.4, 0.2), gN = vec3(0.0, 0.0, 1.0), gR = vec3(1.0, 0.8, 0.0);
vec3 rA = vec3(0.3), rN = vec3(0.0, 1.0, 0.0), rR = vec3(1.0, 0.8, 0.0);
vec3 sA = vec3(0.86, 0.88, 0.92), sN = vec3(0.0, 0.0, 1.0), sR = vec3(1.0, 0.55, 0.0);
vec3 pA, pN, pR;
vec2 uvg = p.xz * 0.28;
vec2 duvgx = dpx.xz * 0.28, duvgy = dpy.xz * 0.28;
// the grass carries the path too (the track is worn into it), and the forest duff
float needGrass = max(grassW, pathW);
float needRock = max(rockW, screeC);
float needSnow = max(snowW, max(snowC, gullyGate));
#ifdef TFAST_3
needGrass = 0.0; needRock = 0.0;
#endif
if (needGrass > 0.002) {
// dry / lush variation across the meadow (read only here and by the carpet below)
float lush = fbmC(cheap, p.xz * 0.006 + 2.0, 2) * 0.5 + 0.5;
matC(cheap, u_grass_d, u_grass_n, u_grass_a, uvg, duvgx, duvgy, gA, gN, gR);
// tint the grass by lushness
gA *= mix(vec3(0.42, 0.55, 0.3), vec3(0.28, 0.55, 0.25), lush) * 0.5;
// sun-facing slopes (south, +z) dry out lighter and warmer; shaded faces stay deep green
gA *= mix(vec3(0.85, 0.92, 0.9), vec3(1.12, 1.06, 0.82), smoothstep(-0.35, 0.35, N.z));
// beyond the blade rings the ground itself carries the clumps: the same cards, seen from above
// Also under the near blades, at reduced weight: the ground seen between standing
// blades must carry the same hue as the carpet that replaces them further out, or
// the field turns from grey-beige to green along a line that walks with the viewer.
float carpetW = mix(0.55, 1.0, smoothstep(8.0, 45.0, dist)) * smoothstep(0.7, 0.35, slope) * grassW;
#ifdef TFAST_3
carpetW = 0.0;
#endif
if (u_carpet_on > 0.5 && carpetW > 0.002) {
vec4 cp = carpetC(cheap, p.xz / 6.0, dpx.xz / 6.0, dpy.xz / 6.0);
// the standing blades in front of it are self-shaded: the carpet is held darker to match them
vec3 cc = cp.rgb * vec3(0.5, 0.57, 0.45) * mix(vec3(0.85, 0.92, 0.9), vec3(1.1, 1.05, 0.85), smoothstep(-0.35, 0.35, N.z)) * (0.75 + 0.35 * lush);
gA = mix(gA, cc, max(cp.a, 0.35) * carpetW * 0.97);
}
// Subalpine forest floor: dark duff where the stands are dense. The photograph-driven
// term marks duff only where the survey actually shows dense conifer, and is the same
// at any distance — ground shading must not depend on where the viewer is.
gA = mix(gA, vec3(0.045, 0.06, 0.025), forestW * 0.85);
}
pA = gA * vec3(0.95, 0.82, 0.62); // the same ground, worn to earth
pN = gN; pR = vec3(0.9, 0.85, 0.0);
// the scanned cliff face on the steep, high slopes; scree below
// The cliff sample and the relief/cliff blend that used to sit here wrote rA/rN/rR
// and were then overwritten wholesale by the rock sample below — they never reached
// the screen (removing them is pixel-identical). Deleting them frees the two sampler
// slots the histogram LUT needs; this shader was at the hardware limit of 16.
if (needRock > 0.002) {
triC(cheap, u_rock_d, u_rock_n, u_rock_a, p, dpx, dpy, N, 0.12, rA, rN, rR);
// macro rock structure for the mountains: a coarse second tile, strata darkening, blue-grey shade side
vec2 uv2 = p.xz * 0.006 + p.y * 0.002;
vec3 rA2 = textureGrad(u_rock_d, uv2, dpx.xz * 0.006 + dpx.y * 0.002, dpy.xz * 0.006 + dpy.y * 0.002).rgb;
vec3 rN2 = textureGrad(u_rock_n, p.zy * 0.01, dpx.zy * 0.01, dpy.zy * 0.01).rgb * 2.0 - 1.0;
rA = mix(rA, rA * rA2 * 2.2, 0.35) * (0.9 + 0.2 * fbmC(cheap, vec2(p.y * 0.03, p.x * 0.004 + p.z * 0.004), 3));
// the Bells' sedimentary strata: near-horizontal bands, tilted a little, sharper on the cliffs
float strata = 0.5 + 0.5 * sin(p.y * 0.45 + p.x * 0.012 + 3.0 * fbmC(cheap, p.xz * 0.01, 2));
rA *= mix(1.0, 0.75 + 0.5 * smoothstep(0.35, 0.65, strata), 0.5 * smoothstep(0.3, 0.6, slope));
rN = normalize(rN + vec3(rN2.x, 0.0, rN2.y) * 0.6 * smoothstep(80.0, 400.0, dist));
// the Bells are maroon mudstone: warm red-brown rock with grey scree below
rA *= mix(vec3(0.14, 0.08, 0.06), vec3(0.24, 0.15, 0.11), fbmC(cheap, p.xz * 0.003, 2) * 0.5 + 0.5);
}
if (needSnow > 0.002) {
sA = textureGrad(u_snow_d, p.xz * 0.25, dpx.xz * 0.25, dpy.xz * 0.25).rgb * 0.8;
}
// ---- blend (height-ish: sharpen with the weights) ----
vec3 alb = gA * grassW + pA * pathW + rA * rockW + sA * snowW;
// the photographed surface (a satellite image of this ground) takes over with distance,
// keeping the scanned materials' fine luminance detail so the middle ground still has grain
float screeMix = 0.0;
if (orthoOn) {
if (screeC > 0.002) {
// near the camera the scanned rocks take the scree at cobble scale (the far talus keeps the coarse tile)
float pebW = smoothstep(220.0, 40.0, dist);
vec3 pebA = vec3(0.36, 0.35, 0.33), pebN = vec3(0.0, 0.0, 1.0);
if (pebW > 0.002) {
pebA = textureGrad(u_rock_d, p.xz * 0.55, dpx.xz * 0.55, dpy.xz * 0.55).rgb * vec3(0.36, 0.35, 0.33);
pebN = textureGrad(u_rock_n, p.xz * 0.55, dpx.xz * 0.55, dpy.xz * 0.55).rgb * 2.0 - 1.0;
}
vec3 screeA = mix(rA * vec3(1.25, 1.2, 1.15), pebA * (0.75 + 0.5 * fbmC(cheap, p.xz * 0.15, 2)), pebW);
// the 10 m photo pixels blur turf and gravel together on the shore: keep grass showing between the cobbles
alb = mix(alb, screeA, screeC * (1.0 - rockW) * mix(0.55, 0.9, smoothstep(30.0, 200.0, dist)));
rN = normalize(mix(rN, normalize(vec3(pebN.x, 1.0, pebN.y)), screeC * pebW * 0.8));
screeMix = screeC;
}
alb = mix(alb, sA, snowC * 0.9);
float lumA = dot(alb, vec3(0.3, 0.59, 0.11));
// How much of the photograph shows through. This was smoothstep(260, 800, dist) —
// ground colour cross-fading toward the survey image as it receded from the camera.
// The photograph has shadows and dark vegetation baked into it from the day it was
// flown, so grass that was plain up close grew dark patches as you backed away, and
// those patches slid and changed shape as you walked. A constant keeps the
// photograph's large-scale colour without tying any of it to the camera.
float orthoW = 0.55;
// grain at three scales so the far slopes keep structure the photograph's pixels cannot carry
#ifdef TFAST_20
float grain = 0.82;
#else
float grain = 0.82 + 0.36 * fbmC(cheap, p.xz * 0.7, 2) + 0.12 * fbmC(cheap, p.xz * 4.0, 2) + 0.3 * (fbmC(cheap, p.xz * 0.06 + 5.0, 3) - 0.5) + 0.15 * (ridgedC(cheap, p.xz * 0.02 + 9.0, 2) - 0.5);
#endif
alb = mix(alb, oc * (0.35 + 1.4 * lumA / max(lumA + 0.12, 1e-3)) * grain, orthoW);
}
// ---- the shoreline, continued onto the land ------------------------------------
// A flat water plane cutting a slope meets it along one exact contour, and no amount
// of shading on the water side removes a mathematically sharp line. The transition
// has to be drawn on BOTH surfaces, so the same wash the water runs is continued up
// the bank here: identical noise fields, identical time, identical phase, keyed off
// height above the lake instead of depth below it. Across the seam the two agree, so
// there is nothing there to read as an edge.
float above = p.y - u_lake_level; // >0 on land, metres
// wet ground: darker and glossier near the water, fading out over ~1.2 m
float wet = smoothstep(1.2, 0.0, above);
alb *= mix(1.0, 0.5, wet * 0.85);
// A sheet of water actually runs up the bank ahead of the foam, so the ground inside
// the wash is seen through water, not bare. Without this the gaps between the foam
// streaks showed dry grass and the wash looked like white paint on a lawn.
float film = smoothstep(0.4, 0.0, above);
vec3 shoreWater = vec3(0.05, 0.11, 0.13) * skyIrradiance(vec3(0, 1, 0)) * 1.15;
alb = mix(alb, mix(alb * 0.5, shoreWater, 0.45), film);
// The wash itself: the water's lap, run above the line and fading as it climbs — the
// same fields, octaves and phase the water uses, so the two agree across the seam.
// Its two gates — the last 22 cm above the waterline, and the first 120 m from the
// camera — are pure geometry, and outside them the four noise fields behind it cannot
// reach the screen. They are worth testing first: the wash is a hairline along one
// shore and the fields were being evaluated for every pixel of the valley.
float washGate = smoothstep(0.22, 0.0, above) * smoothstep(120.0, 15.0, dist);
if (washGate > 0.002) {
float lapT = 0.5 + 0.5 * sin(-above * 9.0 - u_time * 1.6 + 2.0 * gnoiseC(cheap, p.xz * 0.8 + u_time * 0.2));
float fdetT = fbmC(cheap, p.xz * 7.0 - u_time * 0.35, 3) * 0.5 + 0.5;
float fmidT = fbmC(cheap, p.xz * 2.6 + u_time * 0.5, 3) * 0.5 + 0.5;
float fedgeT = fbmC(cheap, p.xz * 1.4 - u_time * 0.3, 2) * 0.5 + 0.5;
// a still alpine lake has a wet line, not surf: keep the wash thin and faint
float fringe = washGate
* (0.12 * smoothstep(0.30, 0.72, fmidT)
+ 0.10 * smoothstep(0.55, 0.95, lapT) * smoothstep(0.22, 0.6, fedgeT)) * (0.55 + 0.75 * fdetT);
alb = mix(alb, vec3(0.72, 0.76, 0.76), clamp(fringe, 0.0, 1.0));
}
// snow lingering in the high gullies (the July photograph's white streaks)
float gully = 0.0;
if (gullyGate > 0.002) { gully = smoothstep(0.62, 0.85, ridgedC(cheap, p.xz * 0.02 + 3.0, 2)) * gullyGate; }
alb = mix(alb, sA * 1.05, gully * 0.9);
vec3 arm = gR * grassW + pR * pathW + rR * rockW + sR * snowW;
// world tangent frame for the planar maps
vec3 T = normalize(vec3(1.0, 0.0, 0.0) - N * N.x);
vec3 B = cross(N, T);
vec3 tn = normalize(gN * grassW + pN * pathW + sN * snowW + vec3(0, 0, 1e-3));
vec3 nPlanar = normalize(T * tn.x + B * tn.y + N * tn.z);
vec3 n = normalize(mix(nPlanar, rN, max(rockW, screeMix * 0.6)));
// fade the detail normal with distance so the far terrain does not sparkle
n = normalize(mix(n, N, smoothstep(150.0, 900.0, dist)));
float ao = arm.r;
float rough = clamp(arm.g, 0.3, 1.0);
float metal = 0.0;
// the shadow map carries the objects standing on the ground (trees, rocks); the
// ground's own relief is in the baked height-field shadow that sunShadow() applies
#ifdef TFAST_1
float shadow = 1.0;
#else
// tersun.frag computed this for exactly this pixel; see the note there for why the
// cascade read cannot happen in here.
float shadow = texelFetch(u_sunshadow, ivec2(gl_FragCoord.xy), 0).r;
#endif
vec3 col = shade(p, n, alb, rough, metal, ao, shadow, viewDepth);
dbg_n = n; dbg_alb = alb; dbg_shadow = shadow; dbg_mat = vec3(rockW, grassW, snowW);
return col;
}
uniform float u_far_split;
uniform float u_far_band;
void main() {
vec3 p = v_wpos;
if (p.y < u_clip_y) discard;
vec3 N = terrainNormal(v_huv);
float slope = 1.0 - N.y;
float dist = length(p - u_cam_pos);
float viewDepth = -(u_view * vec4(p, 1.0)).z;
vec3 col;
#ifdef NEAR_ONLY
// The near program: this patch lies entirely inside the split, so only the detailed
// tier can run here. Compiled alone it does not have to hold the cheap tier's code
// beside it, which is what pushed the combined shader past the register budget.
col = groundShade(p, N, slope, dist, viewDepth, false);
#elif defined(FAR_ONLY)
// The far program: this patch is entirely beyond u_far_split + u_far_band, so every
// pixel in it would take the cheap tier anyway. Compiling that tier on its own — with
// no near path inlined beside it — is the whole point: the two tiers together put this
// shader over the register budget, and the far pixels (most of the screen: the valley
// walls and the Bells) were paying for a near path they never ran.
col = groundShade(p, N, slope, dist, viewDepth, true);
#elif defined(TFAST_5)
col = groundShade(p, N, slope, dist, viewDepth, false);
#else
float band = u_far_band;
if (dist > u_far_split + band) col = groundShade(p, N, slope, dist, viewDepth, true);
else if (dist < u_far_split - band) col = groundShade(p, N, slope, dist, viewDepth, false);
else col = mix(groundShade(p, N, slope, dist, viewDepth, false), groundShade(p, N, slope, dist, viewDepth, true), smoothstep(u_far_split - band, u_far_split + band, dist));
#endif
col = applyFog(col, p, dist);
#ifdef DEBUG_SHADOW
col = vec3(dbg_shadow);
#endif
#ifdef DEBUG_NRM
col = dbg_n * 0.5 + 0.5;
#endif
#ifdef DEBUG_MAT
col = dbg_mat;
#endif
#ifdef DEBUG_ALB
col = dbg_alb * 3.0;
#endif
o_color = vec4(sane(col), 1.0);
}

View file

@ -0,0 +1,50 @@
// CDLOD terrain (Strugar 2009): every draw is one 32x32 patch of the quadtree, placed
// and scaled by u_node. Toward the outer edge of its level's range each vertex morphs
// onto the parent level's grid (odd vertices slide to their even neighbours), so a patch
// meets its coarser neighbour edge-for-edge with no cracks and no popping. Height comes
// from one B-spline sample of the height field at the morphed position, so every level
// sits on the same continuous surface.
layout(location = 0) in vec2 a_xz; // 0..1 across the patch
uniform sampler2D u_height;
uniform float u_half;
uniform mat4 u_view;
uniform mat4 u_proj;
uniform vec2 u_origin;
uniform vec3 u_cam_pos;
uniform vec3 u_node; // x0, z0, size (m)
uniform vec2 u_morph; // distance where the morph starts, and where it is complete
uniform float u_grid; // cells per patch side
float heightSmooth(sampler2D tex, vec2 uv) {
vec2 res = vec2(textureSize(tex, 0));
vec2 t = uv * res - 0.5;
vec2 f = fract(t);
vec2 i = floor(t);
vec2 w0 = (1.0 - f) * (1.0 - f) * (1.0 - f) / 6.0;
vec2 w1 = (4.0 - 6.0 * f * f + 3.0 * f * f * f) / 6.0;
vec2 w3 = f * f * f / 6.0;
vec2 w2 = 1.0 - w0 - w1 - w3;
vec2 s0 = w0 + w1, s1 = w2 + w3;
vec2 o0 = (i - 1.0 + w1 / s0 + 0.5) / res;
vec2 o1 = (i + 1.0 + w3 / s1 + 0.5) / res;
return (texture(tex, vec2(o0.x, o0.y)).r * s0.x + texture(tex, vec2(o1.x, o0.y)).r * s1.x) * s0.y
+ (texture(tex, vec2(o0.x, o1.y)).r * s0.x + texture(tex, vec2(o1.x, o1.y)).r * s1.x) * s1.y;
}
out vec3 v_wpos;
out vec2 v_huv;
void main() {
vec2 grid = a_xz * u_grid;
vec2 xz = u_node.xy + a_xz * u_node.z;
vec2 huv = (xz - u_origin) / (2.0 * u_half) + 0.5;
float h0 = texture(u_height, huv).r;
float d = distance(vec3(xz.x, h0, xz.y), u_cam_pos);
float k = clamp((d - u_morph.x) / max(u_morph.y - u_morph.x, 1.0), 0.0, 1.0);
vec2 frac2 = fract(grid * 0.5) * 2.0; // 1 on odd vertices
grid -= frac2 * k;
xz = u_node.xy + grid / u_grid * u_node.z;
huv = (xz - u_origin) / (2.0 * u_half) + 0.5;
float h = heightSmooth(u_height, huv);
vec3 p = vec3(xz.x, h, xz.y);
v_wpos = p;
v_huv = huv;
gl_Position = u_proj * u_view * vec4(p, 1.0);
}

View file

@ -0,0 +1,39 @@
// Height-field sun shadow, baked once per sun direction (the sun is fixed per scene).
//
// For every height-map texel: the lowest height at which a point above that texel still
// sees the sun. A point (xz, y) is lit iff the ray toward the sun clears the ground
// everywhere along it, i.e. y > h(xz + d.xz t) - d.y t for all t — so the value stored
// is the maximum of that expression over the ray. Every receiver in the scene (ground,
// trunk, crown, card, water) compares its own height against it: one march per texel,
// once, instead of 28 taps per terrain pixel per frame, and vegetation standing in a
// hillside's shadow goes dark with the hillside instead of glowing in front of it.
// The second channel is the distance to the occluder that set the bound, which widens
// the penumbra the way a real shadow softens with distance from its caster.
in vec2 v_uv;
out vec4 o;
uniform sampler2D u_height;
uniform float u_half;
uniform vec3 u_sun;
void main() {
vec2 xz = (v_uv - 0.5) * 2.0 * u_half;
vec3 d = u_sun;
float lit = -1.0e6;
float at = 0.0;
if (d.y > 0.02) {
float t = 1.5, step = 1.5;
for (int i = 0; i < 128; i++) {
vec2 q = xz + d.xz * t;
vec2 uv = q / (2.0 * u_half) + 0.5;
if (uv.x < 0.0 || uv.x > 1.0 || uv.y < 0.0 || uv.y > 1.0) break;
float h = texture(u_height, uv).r - d.y * t;
if (h > lit) { lit = h; at = t; }
t += step;
step *= 1.045;
}
}
// the cloud layer's mask, once, into B: sampled by cloudShadow() with the sun offset and
// the drift applied as a uv shift, instead of a five-octave fbm in every lit pixel of
// every pass
float cloud = smoothstep(0.02, 0.32, fbm(xz * 0.0011, 5));
o = vec4(lit, at, cloud, 1.0);
}

View file

@ -0,0 +1,31 @@
// tersun.frag — the terrain's sun visibility, on its own, one screen-sized R8 buffer.
//
// The ground's shading shader is large: it blends four scanned materials, a photograph
// and a dozen noise fields. Adding a read of the cascade shadow map to it costs about
// six milliseconds a frame on this driver — and costs the same whether the map is tapped
// once or eight times, filtered or texelFetched, compared in hardware or by hand. It is
// a cliff the big shader falls off, not work it performs. The same read from a small
// shader is nearly free, so the read happens here instead: this pass rasterises the same
// CDLOD patches, evaluates the cascades once per pixel, and writes the answer for
// terrain.frag to look up by fragment coordinate.
in vec3 v_wpos;
in vec2 v_huv;
out float o_sh;
uniform sampler2D u_height;
uniform mat4 u_view;
uniform float u_far_split;
uniform float u_far_band;
void main() {
vec3 p = v_wpos;
if (p.y < u_clip_y) discard;
vec3 N = normalize(texture(u_height, v_huv).gba);
float dist = length(p - u_cam_pos);
float viewDepth = -(u_view * vec4(p, 1.0)).z;
// The same tier choice the ground makes, cross-faded over the same band: the near tier
// keeps its filtered penumbra, the far tier its single tap, and the boundary between
// them is not a contour you can find on the hillside.
if (dist > u_far_split + u_far_band) o_sh = sunShadowCheap(p, N, viewDepth);
else if (dist < u_far_split - u_far_band) o_sh = sunShadow(p, N, viewDepth);
else o_sh = mix(sunShadow(p, N, viewDepth), sunShadowCheap(p, N, viewDepth),
smoothstep(u_far_split - u_far_band, u_far_split + u_far_band, dist));
}

View file

@ -0,0 +1,47 @@
// exposure -> ACES -> vignette -> sRGB, with dithering
in vec2 v_uv;
out vec4 o_color;
uniform sampler2D u_hdr;
uniform sampler2D u_bloom;
uniform sampler2D u_ao;
uniform float u_ao_strength;
uniform float u_gi_strength;
uniform vec3 u_wb; // white balance multiplier
uniform vec3 u_lift;
uniform vec3 u_gain;
uniform float u_exposure;
uniform sampler2D u_adapt; // the GPU's adapted exposure (adapt.frag), 1x1
uniform float u_auto; // 1: use it, 0: u_exposure as set
uniform float u_bloom_strength;
uniform float u_vignette;
uniform float u_saturation;
uniform float u_contrast;
vec3 aces(vec3 x) {
const float a = 2.51, b = 0.03, c = 2.43, d = 0.59, e = 0.14;
return clamp((x * (a * x + b)) / (x * (c * x + d) + e), 0.0, 1.0);
}
float hash(vec2 p) { return fract(sin(dot(p, vec2(12.9898, 78.233))) * 43758.5453); }
void main() {
vec3 hdr = sane(texture(u_hdr, v_uv).rgb);
vec4 gi = texture(u_ao, v_uv);
hdr *= mix(1.0, gi.a, u_ao_strength);
// the indirect bounce arrives in the surface's own hue (no albedo buffer in a forward renderer)
float l = dot(hdr, vec3(0.2126, 0.7152, 0.0722));
hdr += gi.rgb * (hdr / max(l, 1e-3)) * u_gi_strength;
vec3 bloom = texture(u_bloom, v_uv).rgb;
float exposure = mix(u_exposure, texture(u_adapt, vec2(0.5)).r, u_auto);
vec3 c = (hdr + bloom * u_bloom_strength) * exposure * u_wb;
// filmic contrast around mid grey in log space
c = max(c, vec3(0.0));
c = pow(c / 0.18, vec3(u_contrast)) * 0.18;
c = aces(c);
// lift / gain grade in display space
c = c * u_gain + u_lift * (1.0 - c);
float lum = dot(c, vec3(0.2126, 0.7152, 0.0722));
c = mix(vec3(lum), c, u_saturation);
vec2 q = v_uv * 2.0 - 1.0;
c *= 1.0 - u_vignette * dot(q, q) * 0.5;
c = pow(c, vec3(1.0 / 2.2));
c += (hash(gl_FragCoord.xy) - 0.5) / 255.0;
o_color = vec4(c, 1.0);
}

View file

@ -0,0 +1,107 @@
// still water: sky reflection with fresnel, sun glitter, scrolling ripple normals, absorption colour
in vec3 v_wpos;
out vec4 o_color;
uniform mat4 u_view;
uniform sampler2D u_depth; // scene depth (resolved) for shore softness / depth tint
uniform mat4 u_inv_vp;
uniform vec2 u_screen;
uniform sampler2D u_refl; // the world mirrored in the surface (rendered by the reflection pass)
uniform float u_refl_on;
uniform sampler2D u_scene; // the scene as drawn before the water: the bed, to refract
// wind-streaked capillary ripples (stretched along the wind) over slower swells
float waterH(vec2 p, float t) {
vec2 w = vec2(p.x * 0.7 + p.y * 0.15, p.y * 1.4) ; // mildly anisotropic: cat's-paws stretched along the wind
// calmer water: the swell keeps most of its weight, the two ripple octaves are
// pulled well down so the surface reads as a lake rather than a chop
return 0.4 * gnoise(w * 0.9 + vec2(t * 0.06, t * 0.4)) + 0.16 * gnoise(p * 2.3 - vec2(t * 0.05, -t * 0.07)) + 0.07 * gnoise(p * 6.0 + vec2(t * 0.9, t * 0.3));
}
vec3 rippleNormal(vec2 p, float t) {
float e = 0.06;
float h = waterH(p, t), hx = waterH(p + vec2(e, 0), t), hz = waterH(p + vec2(0, e), t);
return normalize(vec3(-(hx - h) * 0.26 / e, 1.0, -(hz - h) * 0.26 / e));
}
void main() {
vec3 v = normalize(u_cam_pos - v_wpos);
float dist = length(u_cam_pos - v_wpos);
vec3 n = rippleNormal(v_wpos.xz, u_time);
n = normalize(mix(n, vec3(0, 1, 0), smoothstep(100.0, 600.0, dist))); // calm at a distance
// how deep the ground is under this pixel: from the scene depth
vec2 suv = gl_FragCoord.xy / u_screen;
float sd = texture(u_depth, suv).r;
vec4 gp = u_inv_vp * vec4(suv * 2.0 - 1.0, sd * 2.0 - 1.0, 1.0);
vec3 ground = gp.xyz / gp.w;
float depthBelow = clamp(v_wpos.y - ground.y, 0.0, 10.0);
// How opaque the water is at the shoreline. This used to fade over the last 1.2 m of
// depth, which is the same band the foam lives in, so the surface went transparent
// exactly where it should have been breaking white: the foam was drawn and then
// alpha'd away, leaving a gap of dark wet ground and water that looked like it
// stopped short of the bank. Fade over a much shorter distance so the water reaches
// the edge, and let the foam carry its own opacity below.
vec3 r = reflect(-v, n);
r.y = abs(r.y);
vec3 refl = skyPrefiltered(r, 0.12);
if (u_refl_on > 0.5) {
// the mirrored render lines up with the screen; the ripples nudge and soften the lookup
vec2 ruv = suv + n.xz * 0.02 * smoothstep(500.0, 20.0, dist);
float blur = mix(0.5, 0.2, smoothstep(0.0, 300.0, dist));
refl = sane(textureLod(u_refl, clamp(ruv, 0.001, 0.999), blur).rgb);
}
// wind-blown foam streaks and shoreline wash
float foam = smoothstep(0.62, 0.9, gnoise(vec2(v_wpos.x * 0.25 + u_time * 0.3, v_wpos.z * 1.5) ) * 0.5 + 0.5) * 0.03 * smoothstep(200.0, 30.0, dist);
// Wash: the shallows lapping the shore. Built from fbm rather than one gnoise octave —
// a single octave is a blobby lattice that magnifies into visible squares when you
// stand next to it, which is what made the wash read as cartoon cut-outs. Several
// octaves plus a fine breakup term give it structure at every range it is seen from.
float lap = 0.5 + 0.5 * sin(depthBelow * 9.0 - u_time * 1.6 + 2.0 * gnoise(v_wpos.xz * 0.8 + u_time * 0.2));
float fdet = fbm(v_wpos.xz * 7.0 - u_time * 0.35, 3) * 0.5 + 0.5; // fine bubbles
float fmid = fbm(v_wpos.xz * 2.6 + u_time * 0.5, 3) * 0.5 + 0.5;
float fedge = fbm(v_wpos.xz * 1.4 - u_time * 0.3, 2) * 0.5 + 0.5;
// a still alpine lake has a wet line, not surf: the wash is thin (the last 0.35 m of
// depth) and faint, and the terrain runs the same fields at the same strength
foam += smoothstep(0.35, 0.0, depthBelow) * (0.12 * smoothstep(0.30, 0.72, fmid) + 0.10 * smoothstep(0.55, 0.95, lap) * smoothstep(0.22, 0.6, fedge)) * (0.55 + 0.75 * fdet);
// the lap is a near-field detail: from a distance a lake's edge is a line, not a surf
foam *= smoothstep(120.0, 15.0, dist);
// the wash dies where the surface meets the ground, so it cannot end on a hard line
foam *= smoothstep(0.0, 0.5, length(ground - v_wpos));
float NoV = max(dot(n, v), 0.0);
float F = 0.02 + 0.98 * pow(1.0 - NoV, 5.0);
vec3 hv = normalize(v + u_sun_dir);
float NoH = max(dot(n, hv), 0.0);
float glitter = D_GGX(NoH, 0.06) * 0.25;
float viewDepth = -(u_view * vec4(v_wpos, 1.0)).z;
float shadow = sunShadow(v_wpos, vec3(0, 1, 0), viewDepth) * cloudShadow(v_wpos);
// ---- what is under the surface -------------------------------------------------
// The bed is sampled from the scene as it was drawn before the water, nudged by the
// ripple normal (refraction), then attenuated per channel over the path the light
// actually travelled: down through the water and back up to the eye. Red goes first,
// then green, so shallows stay bright and readable and depth turns blue-green and
// dark on its own. This is what makes it a body of water rather than a tinted sheet:
// the ground is seen through it, not behind it.
vec2 ruv2 = clamp(suv + n.xz * 0.03 * smoothstep(0.0, 2.0, depthBelow), 0.001, 0.999);
// never refract something that is actually in front of the surface (the near bank),
// or the grass on the shore smears out over the water
float rd = texture(u_depth, ruv2).r;
vec4 rgp = u_inv_vp * vec4(ruv2 * 2.0 - 1.0, rd * 2.0 - 1.0, 1.0);
vec3 rground = rgp.xyz / rgp.w;
if (rground.y > v_wpos.y) { ruv2 = suv; }
vec3 bed = sane(texture(u_scene, ruv2).rgb);
float pathLen = depthBelow * (1.0 + 1.0 / max(NoV, 0.25));
vec3 absorb = vec3(0.55, 0.24, 0.14); // per metre: red first, then green — a cold blue-teal depth
vec3 trans = exp(-absorb * pathLen);
vec3 tint = vec3(0.030, 0.085, 0.105) * skyIrradiance(vec3(0, 1, 0)) * 1.15; // Maroon Lake: deep, dark blue-green, not turquoise
vec3 through = bed * trans + tint * (1.0 - trans);
// ---- surface -------------------------------------------------------------------
vec3 col = mix(through, refl, clamp(F * 1.1 + 0.05, 0.0, 0.86)) + u_sun_color * glitter * F * shadow;
col = mix(col, vec3(0.7, 0.75, 0.75) * (skyIrradiance(vec3(0, 1, 0)) * 0.5 + u_sun_color * 0.08 * shadow), clamp(foam, 0.0, 1.0));
col = applyFog(col, v_wpos, dist);
// Soft edge measured ALONG THE VIEW RAY, not vertically. Vertical depth collapses to
// zero over a fraction of a pixel when the surface is seen edge-on, which is exactly
// the low, near-the-waterline view where the plane's silhouette turns into a hard
// glassy line. The distance from the surface to the bed along the ray stays a smooth
// quantity at any angle, so the water dissolves into the ground it meets instead.
float alongRay = length(ground - v_wpos);
float soft = smoothstep(0.0, 0.5, alongRay);
col = mix(bed, col, soft);
o_color = vec4(sane(col), 1.0);
}

View file

@ -0,0 +1,12 @@
layout(location = 0) in vec2 a_xz;
uniform mat4 u_view;
uniform mat4 u_proj;
uniform float u_level;
uniform vec2 u_center;
uniform vec2 u_extent;
out vec3 v_wpos;
void main() {
vec3 p = vec3(u_center.x + a_xz.x * 2.0 * u_extent.x, u_level, u_center.y + a_xz.y * 2.0 * u_extent.y); // the grid spans ±0.5
v_wpos = p;
gl_Position = u_proj * u_view * vec4(p, 1.0);
}

View file

@ -0,0 +1,250 @@
# ============================================================================
# shadow.ludic — cascaded shadow maps for the sun: four 2048^2 depth layers,
# each an orthographic light frustum fitted to the bounding sphere of a slice
# of the camera frustum and snapped to its own texel grid (no swimming).
# ============================================================================
const SHADOW_RES: int = 2048
const SHADOW_CASCADES: int = 5
var sh_tex: int = 0
var sh_fbo: int = 0
var sh_vp: words = null # 4 x 16 float bits
var sh_split: words = null # view-space far distance of each cascade
var sh_range: words = null # 4 light-frustum depth extents (metres)
var sh_texel: words = null # 4 shadow texel sizes (metres)
var sh_tmp_proj: words = null
var sh_tmp_vp: words = null
var sh_tmp_inv: words = null
var sh_tmp_view: words = null
var sh_corner: words = null
var sh_cascade: int = 0 # the cascade being rendered (for casters that skip far ones)
function shadow_init() -> void {
sh_tex = gl_texture()
gl_bind_texture(GL_TEXTURE_2D_ARRAY, sh_tex)
gl_tex_image3d(GL_TEXTURE_2D_ARRAY, 0, GL_DEPTH_COMPONENT32F, SHADOW_RES, SHADOW_RES, SHADOW_CASCADES, 0, GL_DEPTH_COMPONENT, GL_FLOAT, null)
gl_tex_parameteri(GL_TEXTURE_2D_ARRAY, GL_TEXTURE_MIN_FILTER, GL_LINEAR)
gl_tex_parameteri(GL_TEXTURE_2D_ARRAY, GL_TEXTURE_MAG_FILTER, GL_LINEAR)
gl_tex_parameteri(GL_TEXTURE_2D_ARRAY, GL_TEXTURE_WRAP_S, GL_CLAMP_TO_BORDER)
gl_tex_parameteri(GL_TEXTURE_2D_ARRAY, GL_TEXTURE_WRAP_T, GL_CLAMP_TO_BORDER)
gl_tex_parameteri(GL_TEXTURE_2D_ARRAY, GL_TEXTURE_COMPARE_MODE, GL_COMPARE_REF_TO_TEXTURE)
gl_tex_parameteri(GL_TEXTURE_2D_ARRAY, GL_TEXTURE_COMPARE_FUNC, GL_LEQUAL)
let border = gl_floats(4)
gl_put(border, 0, 1.0); gl_put(border, 1, 1.0); gl_put(border, 2, 1.0); gl_put(border, 3, 1.0)
gl_tex_parameterfv(GL_TEXTURE_2D_ARRAY, GL_TEXTURE_BORDER_COLOR, border)
free(border)
sh_fbo = gl_framebuffer()
gl_bind_framebuffer(GL_FRAMEBUFFER, sh_fbo)
gl_draw_buffer(GL_NONE)
gl_read_buffer(GL_NONE)
gl_bind_framebuffer(GL_FRAMEBUFFER, 0)
sh_vp = words(16 * SHADOW_CASCADES)
sh_split = words(SHADOW_CASCADES)
sh_range = words(SHADOW_CASCADES)
sh_texel = words(SHADOW_CASCADES)
# the fourth slice keeps a tree-sized texel out to a kilometre; only the massif uses the last
sh_split[0] = fi(16); sh_split[1] = fi(60); sh_split[2] = fi(250); sh_split[3] = fi(1100); sh_split[4] = fi(6000)
sh_tmp_proj = m4_new(); sh_tmp_vp = m4_new(); sh_tmp_inv = m4_new(); sh_tmp_view = m4_new()
sh_corner = words(3)
}
# light view-projection for the camera-frustum slice [near, far]
function shadow_fit(c: int, near: int, far: int) -> void {
# Fit the slice in VIEW space, not world space. The bounding sphere of a frustum
# slice depends only on near/far/fov/aspect — never on where the camera is pointing —
# so computing it here makes the radius a constant per cascade. Doing it in world
# space (as this did) let the radius wobble as the camera turned, which changed the
# texel size, which moved the grid the projection is snapped to, so the whole shadow
# map resampled every frame: that is the crawl and flicker seen while moving.
m4_perspective(sh_tmp_proj, cam_fov, cam_aspect, near, far)
m4_inverse(sh_tmp_inv, sh_tmp_proj) # NDC -> view space
let cview = v3_new(F_ZERO, F_ZERO, F_ZERO)
let corners = words(24)
for i in 0 .. 8 {
var x = f_neg1(); var y = f_neg1(); var z = f_neg1()
if (i & 1) != 0 { x = F_ONE }
if (i & 2) != 0 { y = F_ONE }
if (i & 4) != 0 { z = F_ONE }
let w = m4_xform_point(sh_corner, sh_tmp_inv, x, y, z)
let iw = f_div(F_ONE, w)
corners[i * 3] = f_mul(sh_corner[0], iw); corners[i * 3 + 1] = f_mul(sh_corner[1], iw); corners[i * 3 + 2] = f_mul(sh_corner[2], iw)
cview[0] = f_add(cview[0], corners[i * 3]); cview[1] = f_add(cview[1], corners[i * 3 + 1]); cview[2] = f_add(cview[2], corners[i * 3 + 2])
}
v3_scale(cview, cview, fr(1, 8))
var radius = F_ZERO
for i in 0 .. 8 {
v3_set(sh_corner, corners[i * 3], corners[i * 3 + 1], corners[i * 3 + 2])
let d = v3_dist(sh_corner, cview)
if f_gt(d, radius) { radius = d }
}
radius = f_mul(radius, fl(1.05))
# the slice centre back into world space
m4_inverse(sh_tmp_vp, cam_view)
let center = words(3)
m4_xform_point(center, sh_tmp_vp, cview[0], cview[1], cview[2])
free(cview)
# light view: from far along the sun direction, looking at the centre
let eye = words(3)
# casters up to ~900 m toward the sun (a mountain across the valley), and the
# slice itself behind the centre: a tight depth range keeps the bias small
let back = f_add(radius, fi(900))
v3_madd(eye, center, sun_dir, back)
let up = v3_new(F_ZERO, F_ONE, F_ZERO)
m4_look_at(sh_tmp_view, eye, center, up)
# snap the ortho window to the shadow texel grid
let texel = f_div(f_mul(radius, F_TWO), fi(SHADOW_RES))
m4_xform_point(sh_corner, sh_tmp_view, center[0], center[1], center[2])
let ox = f_sub(f_mul(f_floor(f_div(sh_corner[0], texel)), texel), sh_corner[0])
let oy = f_sub(f_mul(f_floor(f_div(sh_corner[1], texel)), texel), sh_corner[1])
let nr = f_neg(radius)
let zfar = f_add(f_add(back, radius), fi(100))
m4_ortho(sh_tmp_proj, f_add(nr, ox), f_add(radius, ox), f_add(nr, oy), f_add(radius, oy), F_ONE, zfar)
sh_range[c] = f_sub(zfar, F_ONE)
sh_texel[c] = texel
let out = words(16)
m4_mul(out, sh_tmp_proj, sh_tmp_view)
for i in 0 .. 16 { sh_vp[c * 16 + i] = out[i] }
free(out); free(eye); free(up); free(center); free(corners)
}
function shadow_cascade_vp(c: int) -> words { return mem_off(sh_vp, c * 64) }
# render every cascade; `draw` happens through terrain_draw_shadow + the scene's casters
function shadow_pass() -> void {
var near = cam_near
gl_bind_framebuffer(GL_FRAMEBUFFER, sh_fbo)
gl_viewport(0, 0, SHADOW_RES, SHADOW_RES)
gl_enable(GL_DEPTH_TEST)
gl_depth_func(GL_LESS)
gl_enable(GL_POLYGON_OFFSET_FILL)
gl_polygon_offset(2.0, 4.0)
gl_disable(GL_CULL_FACE)
for c in 0 .. SHADOW_CASCADES {
sh_cascade = c
shadow_fit(c, near, sh_split[c])
near = sh_split[c]
gl_framebuffer_texture_layer(GL_FRAMEBUFFER, GL_DEPTH_ATTACHMENT, sh_tex, 0, c)
if r3d_debug and c == 0 { let st = gl_check_framebuffer_status(GL_FRAMEBUFFER); print(`shadow fbo status {st}`) }
gl_clear(GL_DEPTH_BUFFER_BIT)
let vp = shadow_cascade_vp(c)
# shadows off (a video setting): the cascades stay cleared, so everything reads lit
if sh_enabled {
if not sh_skip_terrain { terrain_draw_shadow(vp) }
scene_draw_casters(vp)
}
}
gl_disable(GL_POLYGON_OFFSET_FILL)
gl_bind_framebuffer(GL_FRAMEBUFFER, 0)
if r3d_debug_shadow { shadow_dump() }
if r3d_debug_shadow and not sh_printed2 {
sh_printed2 = true
let q = words(3)
if sh_probe_x != 0 {
let vp = shadow_cascade_vp(2)
m4_xform_point(q, vp, sh_probe_x, sh_probe_y, sh_probe_z)
print(`probe base ndc {f_fx(q[0])} {f_fx(q[1])} {f_fx(q[2])}`)
m4_xform_point(q, vp, sh_probe_x, f_add(sh_probe_y, fi(15)), sh_probe_z)
print(`probe top ndc {f_fx(q[0])} {f_fx(q[1])} {f_fx(q[2])} -> map texel {f_to_int(f_mul(f_add(f_mul(q[0], F_HALF), F_HALF), fi(SHADOW_RES)))} {f_to_int(f_mul(f_add(f_mul(q[1], F_HALF), F_HALF), fi(SHADOW_RES)))}`)
# where the top's shadow lands on the ground: walk down the sun ray
let gx = f_sub(f_add(sh_probe_x, F_ZERO), f_mul(sun_dir[0], f_div(fi(15), sun_dir[1])))
let gz = f_sub(sh_probe_z, f_mul(sun_dir[2], f_div(fi(15), sun_dir[1])))
m4_xform_point(q, vp, gx, terrain_height(gx, gz), gz)
print(`shadow-of-top ground ndc {f_fx(q[0])} {f_fx(q[1])} {f_fx(q[2])} at {f_fx(gx)} {f_fx(gz)}`)
}
for c in 0 .. SHADOW_CASCADES {
let vp = shadow_cascade_vp(c)
# a point 5 m ahead of the camera on the ground
let px = f_add(cam_pos[0], f_mul(cam_fwd[0], fi(5))); let pz = f_add(cam_pos[2], f_mul(cam_fwd[2], fi(5)))
let w = m4_xform_point(q, vp, px, terrain_height(px, pz), pz)
print(`cascade {c}: ndc {f_fx(q[0])} {f_fx(q[1])} {f_fx(q[2])} w {f_fx(w)} m0 {f_fx(vp[0])} m5 {f_fx(vp[5])} m14 {f_fx(vp[14])}`)
}
free(q)
}
}
var sh_printed2: bool = false
var sh_printed3: bool = false
var sh_enabled: bool = true
var sh_force: int = -1 # R3D_FORCE=<c> pins every pixel to cascade c (debug)
var sh_skip_terrain: bool = false
var sh_probe_x: int = 0
var sh_probe_y: int = 0
var sh_probe_z: int = 0
# Debug: cascade depths as grey PPMs (build/dbg_shadow_<c>.ppm)
function shadow_dump() -> void {
let n = SHADOW_RES * SHADOW_RES
let buf = words(n * SHADOW_CASCADES)
gl_bind_texture(GL_TEXTURE_2D_ARRAY, sh_tex)
gl_tex_parameteri(GL_TEXTURE_2D_ARRAY, GL_TEXTURE_COMPARE_MODE, GL_NONE)
gl_get_tex_image(GL_TEXTURE_2D_ARRAY, 0, GL_DEPTH_COMPONENT, GL_FLOAT, buf)
gl_tex_parameteri(GL_TEXTURE_2D_ARRAY, GL_TEXTURE_COMPARE_MODE, GL_COMPARE_REF_TO_TEXTURE)
if sh_probe_x != 0 {
let vp = shadow_cascade_vp(2)
let q = words(3)
m4_xform_point(q, vp, sh_probe_x, f_add(sh_probe_y, fi(12)), sh_probe_z)
let tx = f_to_int(f_mul(f_add(f_mul(q[0], F_HALF), F_HALF), fi(SHADOW_RES)))
let ty = f_to_int(f_mul(f_add(f_mul(q[1], F_HALF), F_HALF), fi(SHADOW_RES)))
let want = f_add(f_mul(q[2], F_HALF), F_HALF)
print(`probe (12 m up) texel {tx} {ty} card depth {f_fx(f_mul(want, fi(1000)))}/1000`)
for dy in 0 .. 5 {
let yy = ty - 40 + dy * 20
print(` row {yy}: {f_fx(f_mul(buf[2 * n + yy * SHADOW_RES + tx - 20], fi(1000)))} {f_fx(f_mul(buf[2 * n + yy * SHADOW_RES + tx], fi(1000)))} {f_fx(f_mul(buf[2 * n + yy * SHADOW_RES + tx + 20], fi(1000)))} /1000`)
}
free(q)
}
let sm = 512
let row = bytes(sm * 3)
for c in 0 .. SHADOW_CASCADES {
# stretch between the map's own min and max (ignoring the far plane)
var lo = F_ONE; var hi = F_ZERO
var i = 0
while i < n { let d = buf[c * n + i]; if f_ls(d, fl(0.999)) { if f_ls(d, lo) { lo = d }; if f_gt(d, hi) { hi = d } }; i += 97 }
print(`cascade {c} depth range {f_fx(lo)} .. {f_fx(hi)}`)
let f = file_open(`build/dbg_shadow_{c}.ppm`, "wb")
let hdr = `P6\n{sm} {sm}\n255\n`
file_write(f, hdr, len(hdr))
let st = SHADOW_RES / sm
for y in 0 .. sm {
for x in 0 .. sm {
let d = buf[c * n + (y * st) * SHADOW_RES + x * st]
let g = f_to_int(f_mul(f_clamp(f_div(f_sub(d, lo), f_max(f_sub(hi, lo), fl(0.0001))), F_ZERO, F_ONE), fi(255)))
row[x * 3] = g; row[x * 3 + 1] = g; row[x * 3 + 2] = g
}
file_write(f, row, sm * 3)
}
file_close(f)
}
free(buf); free(row)
}
var sh_printed: bool = false
# a uniform array's location: some drivers only answer to the "[0]" spelling
function sh_loc(prog: int, name: string) -> int {
var loc = gl_uniform(prog, name + "[0]")
if loc < 0 { loc = gl_uniform(prog, name) }
return loc
}
function shadow_bind(prog: int) -> void {
r3d_bind_tex(prog, "u_shadow", 15, GL_TEXTURE_2D_ARRAY, sh_tex)
# the height-field shadow (terrain.ludic); a stand-in texture keeps the unit valid before the bake
var ts = ter_shadow_tex
var ts_on = F_ONE
if ts == 0 { ts = ter_height_tex; ts_on = F_ZERO }
r3d_bind_2d(prog, "u_tershadow", 6, ts)
terrain_bind_height(prog)
u_f(gl_uniform(prog, "u_ts_on"), ts_on)
u_f(gl_uniform(prog, "u_ts_half"), fi(TERRAIN_HALF))
u_f2(gl_uniform(prog, "u_ts_origin"), ter_ox, ter_oz)
var loc = gl_uniform(prog, "u_cascade_vp[0]")
if loc < 0 { loc = gl_uniform(prog, "u_cascade_vp") }
if r3d_debug_shadow and not sh_printed { sh_printed = true; print(`cascade vp loc {loc} / {gl_uniform(prog, "u_cascade_vp")} split loc {gl_uniform(prog, "u_cascade_split")} shadow loc {gl_uniform(prog, "u_shadow")}`) }
gl_uniform_matrix4fv(loc, SHADOW_CASCADES, 0, sh_vp)
gl_uniform1fv(sh_loc(prog, "u_cascade_split"), SHADOW_CASCADES, sh_split)
if r3d_debug_shadow and not sh_printed3 { sh_printed3 = true; print(`range {f_fx(sh_range[0])} {f_fx(sh_range[1])} {f_fx(sh_range[2])} {f_fx(sh_range[3])} texel*1000 {f_fx(f_mul(sh_texel[0], fi(1000)))} {f_fx(f_mul(sh_texel[1], fi(1000)))} {f_fx(f_mul(sh_texel[2], fi(1000)))} {f_fx(f_mul(sh_texel[3], fi(1000)))} locs {gl_uniform(prog, "u_cascade_range")} {gl_uniform(prog, "u_cascade_texel")}`) }
gl_uniform1fv(sh_loc(prog, "u_cascade_range"), SHADOW_CASCADES, sh_range)
if Os.has_env("R3D_FORCE") { sh_force = Text.to_int(Os.env("R3D_FORCE")) }
gl_uniform1i(gl_uniform(prog, "u_force_cascade"), sh_force)
gl_uniform1fv(sh_loc(prog, "u_cascade_texel"), SHADOW_CASCADES, sh_texel)
}

View file

@ -0,0 +1,226 @@
# ============================================================================
# skin.ludic — skeletal skinning for glTF models. A Skin is the file's node
# hierarchy (rest translation / rotation / scale per node) plus the skin's joint
# list and inverse bind matrices. The game poses it by giving any node an extra
# rotation and offset IN THE MODEL'S FRAME (X right, Y up, -Z forward, whatever
# the bone's own axes happen to be), skin_pose() folds those into the hierarchy
# and produces the joint matrices, and skin.vert blends four of them per vertex.
#
# Posing in the model frame is what makes a procedural gait writable: "swing the
# thigh forward" is a rotation about the model's X axis, not about whichever axis
# the exporter gave the thigh bone. Per node the delta D is brought into the
# parent's rest frame G_p (the parent's global rest rotation): local rotation =
# (G_p^-1 D G_p) * R_rest.
# ============================================================================
const SKIN_MAX_JOINTS: int = 48
property Skin {
n_nodes: int = 0,
par: words, # parent node per node, -1 at a root
walk: words, # the nodes ordered parents-first
rest_t: words, # 3 per node
rest_r: words, # 4 per node (x, y, z, w)
rest_s: words, # 3 per node
rest_g: words, # 4 per node: the global rest rotation
names: []string,
pose_r: words, # 4 per node: the pose rotation, model frame
pose_t: words, # 3 per node: the pose offset, model frame (metres)
gmat: words, # 16 per node: global matrix this pose
n_joints: int = 0,
joints: words, # node index per joint
inv_bind: words, # 16 per joint
bones: words, # 16 per joint: what the vertex shader skins with
tmp_l: words,
tmp_q: words,
tmp_a: words,
tmp_b: words,
tmp_c: words,
tmp_v: words
}
# a 3-vector of a JSON array (float bits), or a default
function skin_jv3(o: words, at: int, nd: Val, key: pointer, dx: int, dy: int, dz: int) -> void {
if value_has(nd, key) == 0 { o[at] = dx; o[at + 1] = dy; o[at + 2] = dz; return }
let arr = value_get(nd, key)
for i in 0 .. 3 { o[at + i] = jnum(value_at(arr, i)) }
}
# JOINTS_0 / WEIGHTS_0 onto attributes 5 and 6 of the VAO being built (gltf_prim)
function skin_attribs(m: Mesh, attrs: Val) -> bool {
if value_has(attrs, "JOINTS_0") == 0 or value_has(attrs, "WEIGHTS_0") == 0 { return false }
let jd = gltf_accessor(value_as_int(value_get(attrs, "JOINTS_0")))
var jsz = 1
var jtype = GL_UNSIGNED_BYTE
if gltf_ctype == 5123 { jsz = 2; jtype = GL_UNSIGNED_SHORT }
let jb = gl_buffer()
gl_bind_buffer(GL_ARRAY_BUFFER, jb)
gl_buffer_data(GL_ARRAY_BUFFER, gltf_count * gltf_comps * jsz, jd, GL_STATIC_DRAW)
gl_enable_vertex_attrib_array(5)
gl_vertex_attrib_pointer(5, gltf_comps, jtype, 0, 0, null) # integers, read as floats
free(jd)
let wd = gltf_accessor(value_as_int(value_get(attrs, "WEIGHTS_0")))
var wsz = 4
var wtype = GL_FLOAT
var norm = 0
if gltf_ctype == 5123 { wsz = 2; wtype = GL_UNSIGNED_SHORT; norm = 1 }
if gltf_ctype == 5121 { wsz = 1; wtype = GL_UNSIGNED_BYTE; norm = 1 }
let wb = gl_buffer()
gl_bind_buffer(GL_ARRAY_BUFFER, wb)
gl_buffer_data(GL_ARRAY_BUFFER, gltf_count * gltf_comps * wsz, wd, GL_STATIC_DRAW)
gl_enable_vertex_attrib_array(6)
gl_vertex_attrib_pointer(6, gltf_comps, wtype, norm, 0, null)
free(wd)
return true
}
# the skin `idx` of the document being loaded (gltf_load holds gltf_doc / gltf_bin open)
function skin_load(idx: int) -> Skin {
let sk = new Skin
let nodes = value_get(gltf_doc, "nodes")
let n = value_count(nodes)
sk.n_nodes = n
sk.par = words(n); sk.walk = words(n)
sk.rest_t = words(n * 3); sk.rest_r = words(n * 4); sk.rest_s = words(n * 3); sk.rest_g = words(n * 4)
sk.pose_r = words(n * 4); sk.pose_t = words(n * 3); sk.gmat = words(n * 16)
sk.names = new []string
sk.tmp_l = m4_new(); sk.tmp_q = q_new(); sk.tmp_a = q_new(); sk.tmp_b = q_new(); sk.tmp_c = q_new(); sk.tmp_v = words(3)
for i in 0 .. n { sk.par[i] = -1 }
for i in 0 .. n {
let nd = value_at(nodes, i)
var nm: string = ""
if value_has(nd, "name") != 0 { nm = value_as_str(value_get(nd, "name")) }
push(sk.names, nm)
skin_jv3(sk.rest_t, i * 3, nd, "translation", F_ZERO, F_ZERO, F_ZERO)
skin_jv3(sk.rest_s, i * 3, nd, "scale", F_ONE, F_ONE, F_ONE)
if value_has(nd, "rotation") != 0 {
let r = value_get(nd, "rotation")
for k in 0 .. 4 { sk.rest_r[i * 4 + k] = jnum(value_at(r, k)) }
} else { sk.rest_r[i * 4] = F_ZERO; sk.rest_r[i * 4 + 1] = F_ZERO; sk.rest_r[i * 4 + 2] = F_ZERO; sk.rest_r[i * 4 + 3] = F_ONE }
if value_has(nd, "matrix") != 0 { print(`skin: node {nm} uses a matrix transform (unsupported, treated as identity)`) }
if value_has(nd, "children") != 0 {
let ch = value_get(nd, "children")
for k in 0 .. value_count(ch) { sk.par[value_as_int(value_at(ch, k))] = i }
}
}
# parents first: order the nodes by depth
let depth = words(n)
for i in 0 .. n {
var d = 0
var p = sk.par[i]
while p >= 0 and d < n { d += 1; p = sk.par[p] }
depth[i] = d
}
var k = 0
for d in 0 .. n { for i in 0 .. n { if depth[i] == d { sk.walk[k] = i; k += 1 } } }
free(depth)
# the global rest rotation of every node
for w in 0 .. n {
let i = sk.walk[w]
let p = sk.par[i]
q_load(sk.tmp_a, sk.rest_r, i)
if p >= 0 { q_load(sk.tmp_b, sk.rest_g, p); q_mul(sk.tmp_q, sk.tmp_b, sk.tmp_a); q_store(sk.rest_g, i, sk.tmp_q) }
else { q_store(sk.rest_g, i, sk.tmp_a) }
}
# the skin: joints and inverse bind matrices
let skv = value_at(value_get(gltf_doc, "skins"), idx)
let jl = value_get(skv, "joints")
var nj = value_count(jl)
if nj > SKIN_MAX_JOINTS { print(`skin: {nj} joints, only the first {SKIN_MAX_JOINTS} are used`); nj = SKIN_MAX_JOINTS }
sk.n_joints = nj
sk.joints = words(nj)
sk.inv_bind = words(nj * 16)
sk.bones = words(nj * 16)
for j in 0 .. nj { sk.joints[j] = value_as_int(value_at(jl, j)) }
if value_has(skv, "inverseBindMatrices") != 0 {
let ib = gltf_accessor(value_as_int(value_get(skv, "inverseBindMatrices")))
for i in 0 .. nj * 16 { sk.inv_bind[i] = mem_get_f32_bits(ib, i) }
free(ib)
} else {
for j in 0 .. nj { m4_identity(mem_off(sk.inv_bind, j * 64)) }
}
skin_reset(sk)
skin_pose(sk)
print(`skin: {nj} joints over {n} nodes`)
return sk
}
function skin_find(sk: Skin, name: string) -> int {
for i in 0 .. sk.n_nodes { if sk.names[i] == name { return i } }
print(`skin: no node {name}`)
return -1
}
function skin_mat(sk: Skin, node: int) -> words { return mem_off(sk.gmat, node * 64) }
# back to the rest pose
function skin_reset(sk: Skin) -> void {
for i in 0 .. sk.n_nodes {
sk.pose_r[i * 4] = F_ZERO; sk.pose_r[i * 4 + 1] = F_ZERO; sk.pose_r[i * 4 + 2] = F_ZERO; sk.pose_r[i * 4 + 3] = F_ONE
sk.pose_t[i * 3] = F_ZERO; sk.pose_t[i * 3 + 1] = F_ZERO; sk.pose_t[i * 3 + 2] = F_ZERO
}
}
# a node's pose rotation in the model frame: pitch about X, yaw about Y, roll about Z (radians)
function skin_set_rot(sk: Skin, node: int, pitch: int, yaw: int, roll: int) -> void {
if node < 0 { return }
q_euler(sk.tmp_q, pitch, yaw, roll)
q_store(sk.pose_r, node, sk.tmp_q)
}
function skin_set_quat(sk: Skin, node: int, q: words) -> void { if node >= 0 { q_store(sk.pose_r, node, q) } }
# a node's pose offset in the model frame (metres)
function skin_set_offset(sk: Skin, node: int, x: int, y: int, z: int) -> void {
if node < 0 { return }
sk.pose_t[node * 3] = x; sk.pose_t[node * 3 + 1] = y; sk.pose_t[node * 3 + 2] = z
}
# fold the pose into the hierarchy: global matrices, then the joint matrices
function skin_pose(sk: Skin) -> void {
for w in 0 .. sk.n_nodes {
let i = sk.walk[w]
let p = sk.par[i]
q_load(sk.tmp_a, sk.pose_r, i) # D, model frame
var tx = sk.rest_t[i * 3]; var ty = sk.rest_t[i * 3 + 1]; var tz = sk.rest_t[i * 3 + 2]
let ox = sk.pose_t[i * 3]; let oy = sk.pose_t[i * 3 + 1]; let oz = sk.pose_t[i * 3 + 2]
if p >= 0 {
q_load(sk.tmp_b, sk.rest_g, p) # G_p
q_conj(sk.tmp_c, sk.tmp_b) # G_p^-1
q_mul(sk.tmp_q, sk.tmp_c, sk.tmp_a)
q_mul(sk.tmp_a, sk.tmp_q, sk.tmp_b) # G_p^-1 D G_p
if ox != 0 or oy != 0 or oz != 0 {
v3_set(sk.tmp_v, ox, oy, oz)
q_rotate(sk.tmp_v, sk.tmp_c, sk.tmp_v) # the offset in the parent's frame
tx = f_add(tx, sk.tmp_v[0]); ty = f_add(ty, sk.tmp_v[1]); tz = f_add(tz, sk.tmp_v[2])
}
} else { tx = f_add(tx, ox); ty = f_add(ty, oy); tz = f_add(tz, oz) }
q_load(sk.tmp_b, sk.rest_r, i)
q_mul(sk.tmp_q, sk.tmp_a, sk.tmp_b) # local rotation
m4_trs_q(sk.tmp_l, tx, ty, tz, sk.tmp_q, sk.rest_s[i * 3], sk.rest_s[i * 3 + 1], sk.rest_s[i * 3 + 2])
if p >= 0 { m4_mul(skin_mat(sk, i), skin_mat(sk, p), sk.tmp_l) }
else { m4_copy(skin_mat(sk, i), sk.tmp_l) }
}
for j in 0 .. sk.n_joints {
m4_mul(mem_off(sk.bones, j * 64), skin_mat(sk, sk.joints[j]), mem_off(sk.inv_bind, j * 64))
}
}
# the joint matrices onto a program's u_bones[]
function skin_bind(sk: Skin, prog: int) -> void {
var loc = gl_uniform(prog, "u_bones[0]")
if loc < 0 { loc = gl_uniform(prog, "u_bones") }
gl_uniform_matrix4fv(loc, sk.n_joints, 0, sk.bones)
}
# the same skeleton posed on its own: shares the rest data, owns the pose and the matrices
function skin_clone(src: Skin) -> Skin {
let sk = new Skin
sk.n_nodes = src.n_nodes; sk.par = src.par; sk.walk = src.walk
sk.rest_t = src.rest_t; sk.rest_r = src.rest_r; sk.rest_s = src.rest_s; sk.rest_g = src.rest_g
sk.names = src.names
sk.n_joints = src.n_joints; sk.joints = src.joints; sk.inv_bind = src.inv_bind
let n = src.n_nodes
sk.pose_r = words(n * 4); sk.pose_t = words(n * 3); sk.gmat = words(n * 16)
sk.bones = words(src.n_joints * 16)
sk.tmp_l = m4_new(); sk.tmp_q = q_new(); sk.tmp_a = q_new(); sk.tmp_b = q_new(); sk.tmp_c = q_new(); sk.tmp_v = words(3)
skin_reset(sk)
skin_pose(sk)
return sk
}

View file

@ -0,0 +1,134 @@
# ============================================================================
# sky.ludic — the HDRI sky and its image-based lighting: the equirect radiance
# map (RGB16F, mipped), the sun found in it, a diffuse-convolved irradiance map,
# a GGX-prefiltered map per roughness level (a 2D array), and the split-sum
# BRDF lookup. All convolved on the GPU at load.
# ============================================================================
const SKY_PREFILTER_LEVELS: int = 6
var sky_tex: int = 0
var sky_w: int = 0
var sky_h: int = 0
var sky_irradiance: int = 0
var sky_prefilter: int = 0 # GL_TEXTURE_2D_ARRAY
var sky_brdf: int = 0
var sun_dir: words = null # toward the sun (float bits)
var sun_color: words = null # radiance (float bits)
var sky_fullscreen: Mesh = null
var sky_yaw: int = 0 # radians: the HDRI is turned by this about y
var sky_sun_boost: int = 0x40133333 # 2.3: the photograph's thin cloud dims its sun; a crisper day wants more
var sky_rot_s: int = 0
var sky_rot_c: int = 0
var sun_hdri: words = null # the sun direction as found in the file
var sky_p_irr: int = 0
var sky_p_pre: int = 0
var sky_p_brdf: int = 0
# turn the HDRI so its sun sits at world azimuth `yaw` (radians, 0 = toward -z)
function sky_set_yaw(yaw: int) -> void {
sky_set_rot(yaw)
# world sun = rotY(sun_hdri, -yaw): the lookup rotates a world direction by +yaw
let s = sun_hdri
v3_set(sun_dir, f_sub(f_mul(sky_rot_c, s[0]), f_mul(sky_rot_s, s[2])), s[1], f_add(f_mul(sky_rot_s, s[0]), f_mul(sky_rot_c, s[2])))
sky_precompute()
}
# turn only the visible sky image (cheap, per frame): the light and the convolved
# maps stay where they are — daylight.ludic moves those on its own terms
function sky_set_rot(yaw: int) -> void {
sky_yaw = yaw
sky_rot_s = f_sin(yaw); sky_rot_c = f_cos(yaw)
}
function sky_bind_rot(prog: int) -> void { u_f2(gl_uniform(prog, "u_sky_rot"), sky_rot_s, sky_rot_c) }
# direction for an equirect uv (matches equirectUV in lighting.glsl)
function sky_dir_from_uv(o: words, u: int, v: int) -> void {
let phi = f_mul(f_sub(u, F_HALF), f_mul(F_TWO, F_PI))
let theta = f_mul(v, F_PI)
let st = f_sin(theta)
v3_set(o, f_mul(st, f_sin(phi)), f_cos(theta), f_neg(f_mul(st, f_cos(phi))))
}
function sky_load(path: string) -> bool {
sky_tex = tex_load_hdr(path)
if sky_tex == 0 { return false }
sky_w = tex_w; sky_h = tex_h
sun_dir = words(3)
sun_hdri = words(3)
sky_dir_from_uv(sun_hdri, fr(hdr_max_x * 2 + 1, sky_w * 2), fr(hdr_max_y * 2 + 1, sky_h * 2))
v3_copy(sun_dir, sun_hdri)
sky_rot_c = F_ONE
# the sun's irradiance is what the IBL clip leaves out of the map; lighting it
# directly with that keeps sun and sky in the photograph's own proportion
sun_color = v3_new(f_mul(hdr_sun_r, sky_sun_boost), f_mul(hdr_sun_g, sky_sun_boost), f_mul(hdr_sun_b, sky_sun_boost))
print(`sun irradiance: {f_fx(hdr_sun_r)} {f_fx(hdr_sun_g)} {f_fx(hdr_sun_b)} (Q16.16), clip {f_fx(hdr_clip)}`)
print(`sky: {sky_w}x{sky_h}, sun at texel {hdr_max_x},{hdr_max_y}`)
sky_fullscreen = mesh_fullscreen()
sky_precompute()
return true
}
function sky_convolve(prog: int, target_tex: int, layer: int, w: int, h: int, rough: int) -> void {
let fbo = gl_framebuffer()
gl_bind_framebuffer(GL_FRAMEBUFFER, fbo)
if layer < 0 { gl_framebuffer_texture2d(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, target_tex, 0) }
else { gl_framebuffer_texture_layer(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, target_tex, 0, layer) }
gl_viewport(0, 0, w, h)
gl_use_program(prog)
r3d_bind_2d(prog, "u_sky", 0, sky_tex)
sky_bind_rot(prog)
u_f(gl_uniform(prog, "u_sun_clip"), hdr_clip)
u_f(gl_uniform(prog, "u_rough"), rough)
u_f(gl_uniform(prog, "u_sky_w"), fi(sky_w))
mesh_draw(sky_fullscreen)
gl_bind_framebuffer(GL_FRAMEBUFFER, 0)
let ids = gl_scratch()
ids[0] = fbo
gl_delete_framebuffers(1, ids)
}
function sky_precompute() -> void {
gl_disable(GL_DEPTH_TEST)
if sky_irradiance != 0 {
let ids = gl_scratch()
ids[0] = sky_irradiance; gl_delete_textures(1, ids)
ids[0] = sky_prefilter; gl_delete_textures(1, ids)
ids[0] = sky_brdf; gl_delete_textures(1, ids)
}
# irradiance: 128 x 64 equirect
if sky_p_irr == 0 { sky_p_irr = r3d_program("fullscreen.vert", "ibl_irradiance.frag", ""); sky_p_pre = r3d_program("fullscreen.vert", "ibl_prefilter.frag", ""); sky_p_brdf = r3d_program("fullscreen.vert", "ibl_brdf.frag", "") }
let p_irr = sky_p_irr
sky_irradiance = tex_target(128, 64, GL_RGB16F, GL_RGB, GL_FLOAT, GL_LINEAR)
gl_bind_texture(GL_TEXTURE_2D, sky_irradiance)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_WRAP_S, GL_REPEAT)
sky_convolve(p_irr, sky_irradiance, -1, 128, 64, F_ZERO)
# prefiltered specular: 6 roughness levels, 512 x 256 each, as a 2D array
let p_pre = sky_p_pre
sky_prefilter = gl_texture()
gl_bind_texture(GL_TEXTURE_2D_ARRAY, sky_prefilter)
gl_tex_image3d(GL_TEXTURE_2D_ARRAY, 0, GL_RGB16F, 512, 256, SKY_PREFILTER_LEVELS, 0, GL_RGB, GL_FLOAT, null)
gl_tex_parameteri(GL_TEXTURE_2D_ARRAY, GL_TEXTURE_WRAP_S, GL_REPEAT)
gl_tex_parameteri(GL_TEXTURE_2D_ARRAY, GL_TEXTURE_WRAP_T, GL_CLAMP_TO_EDGE)
gl_tex_parameteri(GL_TEXTURE_2D_ARRAY, GL_TEXTURE_MAG_FILTER, GL_LINEAR)
gl_tex_parameteri(GL_TEXTURE_2D_ARRAY, GL_TEXTURE_MIN_FILTER, GL_LINEAR)
for l in 0 .. SKY_PREFILTER_LEVELS {
sky_convolve(p_pre, sky_prefilter, l, 512, 256, fr(l, SKY_PREFILTER_LEVELS - 1))
}
# BRDF LUT
let p_brdf = sky_p_brdf
sky_brdf = tex_target(256, 256, GL_RG16F, GL_RG, GL_FLOAT, GL_LINEAR)
sky_convolve(p_brdf, sky_brdf, -1, 256, 256, F_ZERO)
gl_check("sky precompute")
}
# bind the IBL set + sun for a lit program (units 12..14)
function sky_bind_lighting(prog: int) -> void {
r3d_bind_2d(prog, "u_irradiance", 12, sky_irradiance)
r3d_bind_tex(prog, "u_prefilter", 13, GL_TEXTURE_2D_ARRAY, sky_prefilter)
r3d_bind_2d(prog, "u_brdf", 14, sky_brdf)
u_v3(gl_uniform(prog, "u_sun_dir"), sun_dir)
u_v3(gl_uniform(prog, "u_sun_color"), sun_color)
u_v3(gl_uniform(prog, "u_cam_pos"), cam_pos)
u_f(gl_uniform(prog, "u_prefilter_levels"), fi(SKY_PREFILTER_LEVELS))
daylight_bind(prog)
}

View file

@ -0,0 +1,315 @@
# ============================================================================
# stream.ludic — ground cover that follows the camera anywhere on the map.
#
# The world is cut into square chunks. A chunk's instances are generated once
# per distance band (a deterministic function of the chunk and the band, so the
# same ground always grows the same grass) and cached; each frame the chunks
# within reach are gathered into the layer's instance list. Bands thin the cover
# with distance and grow the cards so the carpet stays continuous on screen;
# beyond the last band nothing is placed (the terrain material carries it).
#
# The scene supplies the generator: stream_fill(chunk_x, chunk_z, band) calls
# stream_emit(...) per instance. One Stream drives one Layer.
# ============================================================================
var STREAM_MAX_CHUNKS: int = 4096
property Chunk {
key: int = 0, # packed (cx, cz, band)
used: int = 0, # the walk that last wanted it (for eviction)
data: words, # INST_FLOATS per instance
count: int = 0,
ymin: int = 0, # height range of its instances (float bits), for the frustum test
ymax: int = 0
}
property Stream {
layer: Layer,
size: int = 0, # chunk size (metres, float bits)
reach: int = 0, # radius (metres, float bits)
bands: words, # band outer radii (float bits), ascending; 4 of them
chunks: []Chunk,
keys: words, # parallel to chunks for lookup
n: int = 0,
last_cx: int = 999999,
last_cz: int = 999999,
pending: bool = false, # chunks still to generate after the camera crossed a cell
cur: Chunk, # the chunk being filled
kind: int = 0, # the scene's generator selector for this stream
min_band: int = 0, # bands below this belong to another (nearer) stream
view_gen: int = -1, # sc_view_gen the layer was last gathered for (the view turned -> regather)
htab: words # open-addressed key -> chunk index + 1 (0 = empty)
}
var stream_all: []Stream = null
var stream_cap_read: bool = false
var stream_no_evict: bool = false # R3D_NOEVICT: the old behaviour, for comparison
var stream_walk_no: int = 0 # counts ring walks; a chunk's age is measured in these
var stream_evictions: int = 0
# microseconds spent per frame, split so the hitch can be attributed (R3D_PROF=1)
var stream_us_gen: long = 0 # generating new chunks (stream_fill)
var stream_us_gather: long = 0 # copying cached chunks into the layer buffer
var stream_us_walk: long = 0 # the ring walk itself
var stream_walks: int = 0 # streams that walked their whole ring this frame
function stream_new(layer: Layer, size: int, reach: int, b0: int, b1: int, b2: int, b3: int) -> Stream {
if not stream_cap_read {
stream_cap_read = true
if Os.has_env("R3D_STREAM_CAP") { STREAM_MAX_CHUNKS = Text.to_int(Os.env("R3D_STREAM_CAP")) }
stream_no_evict = Os.has_env("R3D_NOEVICT")
}
let s = new Stream
s.layer = layer; s.size = size; s.reach = reach
layer.streamed = true
layer.grounded = true
s.bands = words(4)
s.bands[0] = b0; s.bands[1] = b1; s.bands[2] = b2; s.bands[3] = b3
s.chunks = new []Chunk
s.keys = words(STREAM_MAX_CHUNKS)
s.htab = words(STREAM_HASH)
for i in 0 .. STREAM_HASH { s.htab[i] = 0 }
if stream_all == null { stream_all = new []Stream }
push(stream_all, s)
return s
}
function stream_key(cx: int, cz: int, band: int) -> int { return ((cx + 4096) * 8192 + (cz + 4096)) * 4 + band }
# Chunk lookup is an open-addressed hash, not a scan. A cell crossing tests every cell
# within reach — for the 800 m streams that is ~2000 cells each — and a scan over the
# cached chunks made that O(cells x chunks), tens of millions of comparisons in the one
# frame that crosses a 32 m boundary, growing as more ground is explored. That is the
# stutter you feel when walking, and it never shows in a stationary profile because
# stream_update returns immediately while the camera stays in its cell.
const STREAM_HASH: int = 8192 # power of two, >= 2 * STREAM_MAX_CHUNKS
function stream_slot(key: int) -> int {
var h = key * -1640531527 # Knuth's golden-ratio multiplier, as a signed i32
h = h ^ (h >> 15)
return h & (STREAM_HASH - 1)
}
function stream_find(s: Stream, key: int) -> Chunk {
var i = stream_slot(key)
while s.htab[i] != 0 {
let idx = s.htab[i] - 1
if s.keys[idx] == key { return s.chunks[idx] }
i = (i + 1) & (STREAM_HASH - 1)
}
return null
}
function stream_remember(s: Stream, key: int, idx: int) -> void {
var i = stream_slot(key)
while s.htab[i] != 0 { i = (i + 1) & (STREAM_HASH - 1) }
s.htab[i] = idx + 1
}
# the generator adds instances to the chunk being filled (into a shared scratch; the
# chunk gets an exactly-sized copy when the fill ends)
const STREAM_CHUNK_MAX: int = 262144
var stream_debug_n: int = 0
var stream_scratch: words = null
function stream_emit(s: Stream, x: int, y: int, z: int, scale: int, yaw: int, seed: int, wind: int) -> void {
let c = s.cur
if c.count >= STREAM_CHUNK_MAX { return }
if stream_scratch == null { stream_scratch = words(STREAM_CHUNK_MAX * INST_FLOATS) }
if c.count == 0 { c.ymin = y; c.ymax = y } else { c.ymin = f_min(c.ymin, y); c.ymax = f_max(c.ymax, y) }
let o = c.count * INST_FLOATS
stream_scratch[o] = x; stream_scratch[o + 1] = y; stream_scratch[o + 2] = z; stream_scratch[o + 3] = scale
stream_scratch[o + 4] = f_sin(yaw); stream_scratch[o + 5] = f_cos(yaw); stream_scratch[o + 6] = seed; stream_scratch[o + 7] = wind
c.count += 1
}
function stream_band(s: Stream, d: int) -> int {
if f_ls(d, s.bands[0]) { return 0 }
if f_ls(d, s.bands[1]) { return 1 }
if f_ls(d, s.bands[2]) { return 2 }
if f_ls(d, s.bands[3]) { return 3 }
return 4
}
# gather the chunks around the camera into the layer, generating missing ones nearest
# first within a per-frame budget so a cell crossing spreads over a few frames instead of
# one hitch (the very first update, before anything is on screen, generates everything)
# The budget is global, not per stream. It used to be 60000 per stream, and with the
# thirteen streams a scene like the valley runs that let a single frame generate over
# 700k instances — so crossing a 32 m cell put one frame's worth of cover generation
# (height samples, ortho lookups, slope and path tests, per candidate) into one frame
# while its neighbours did none. That one frame is the stutter you feel while walking;
# spreading the same work over several frames costs nothing but a little pop-in at the
# far edge of the reach, where new chunks appear.
# A time budget, not an instance count. Instances are a poor proxy: a candidate that
# is rejected costs nearly as much as one that is kept, and cost per instance varies
# by band and kind. With a real microsecond clock the budget can just be the thing we
# actually care about — how long this frame is allowed to spend growing ground cover.
# Overshoot is bounded by one chunk, so keep chunks small on the dense near streams.
const STREAM_BUDGET_US: int = 2500 # microseconds of generation per frame
var stream_deadline: long = 0
const STREAM_BUDGET: int = 8000 # kept for the work counter only
# The worst frame is now bounded by one chunk, not by the budget: stream_fill emits a
# whole chunk in one call, and the densest band-0 chunk is ~114k instances. Splitting a
# chunk's generation across frames would need a resumable generator contract; that is
# the next step if the residual hitch ever matters.
var stream_budget_left: int = 0
# Drop the half of the cache nobody has asked for in the longest time, and rebuild the
# index over what is left.
#
# Before this, a full cache simply stopped remembering: the chunk was generated, used for
# that frame and thrown away, so every walk regenerated it. That is not a slow degradation
# — it is a cliff. Past it every frame pays the whole generation budget and the ground
# visibly re-grows as you turn, and it arrives after enough of the map has been walked,
# which is exactly when a player is least likely to connect it to anything.
function stream_evict(s: Stream) -> void {
# the age threshold that keeps about half, found by bisection on the count (no sort)
var lo = 0
var hi = stream_walk_no
var keep = s.n / 2
var t = 0
var it = 0
while it < 24 and lo < hi {
t = (lo + hi + 1) / 2
var c = 0
var i = 0
while i < s.n { if s.chunks[i].used >= t { c += 1 }; i += 1 }
if c >= keep { lo = t } else { hi = t - 1 }
it += 1
}
t = lo
# everything wanted by the walk in progress stays whatever the threshold says
let kept = new []Chunk
var i = 0
while i < s.n {
let c = s.chunks[i]
if c.used >= t or c.used == stream_walk_no { push(kept, c) }
else { if c.data != null { free(c.data) } }
i += 1
}
s.chunks = kept
s.n = len(kept)
for h in 0 .. STREAM_HASH { s.htab[h] = 0 }
i = 0
while i < s.n { s.keys[i] = s.chunks[i].key; stream_remember(s, s.chunks[i].key, i); i += 1 }
stream_evictions += 1
}
function stream_update(s: Stream, cam_x: int, cam_z: int) -> void {
let ccx = f_to_int(f_floor(f_div(cam_x, s.size)))
let ccz = f_to_int(f_floor(f_div(cam_z, s.size)))
if ccx == s.last_cx and ccz == s.last_cz and not s.pending and s.view_gen == sc_view_gen { return }
let first = s.last_cx == 999999
s.view_gen = sc_view_gen
s.last_cx = ccx; s.last_cz = ccz
let l = s.layer
l.count = 0
var missing = false
stream_walks += 1
stream_walk_no += 1
let tw = gl_now_us()
let r = f_to_int(f_div(s.reach, s.size)) + 1
# rings outward from the camera's cell: the nearest chunks are generated first
var ring = 0
while ring <= r {
var cz = ccz - ring
while cz <= ccz + ring {
var cx = ccx - ring
while cx <= ccx + ring {
let edge = (cz == ccz - ring) or (cz == ccz + ring) or (cx == ccx - ring) or (cx == ccx + ring)
if edge {
let wx = f_mul(f_add(fi(cx), F_HALF), s.size)
let wz = f_mul(f_add(fi(cz), F_HALF), s.size)
let dx = f_sub(wx, cam_x); let dz = f_sub(wz, cam_z)
let d = f_sqrt(f_add(f_mul(dx, dx), f_mul(dz, dz)))
let band = stream_band(s, d)
if band < 4 and band >= s.min_band and f_ls(d, f_add(s.reach, s.size)) {
let key = stream_key(cx, cz, band)
var c = stream_find(s, key)
if c != null { c.used = stream_walk_no }
# The cell underfoot and its neighbours are never deferred: they are what you
# are looking at, and a hole there is the grass vanishing as you walk into it.
let urgent = band == 0 and ring <= 1
if c == null and (first or urgent or gl_now_us() < stream_deadline) {
c = new Chunk
c.key = key
s.cur = c
let t0 = gl_now_us()
stream_fill(s, cx, cz, band)
let dt = gl_now_us() - t0
stream_us_gen = stream_us_gen + dt
prof_chunk(s.kind, band, c.count, dt)
if r3d_debug and band == 0 and stream_debug_n < 40 { stream_debug_n += 1; print(`stream kind {s.kind} band {band} chunk {cx},{cz}: {c.count} instances`) }
if c.count > 0 { c.data = words(c.count * INST_FLOATS); mem_copy(c.data, stream_scratch, c.count * INST_FLOATS * 4) }
if s.n >= STREAM_MAX_CHUNKS and not stream_no_evict { stream_evict(s) }
# If the walk in progress wants more chunks than the cache can hold, there
# is nothing to evict and this one is used and dropped, as every chunk used
# to be. The cap has to exceed one walk's ring for the cache to work at all.
if s.n < STREAM_MAX_CHUNKS {
push(s.chunks, c); s.keys[s.n] = key; stream_remember(s, key, s.n); s.n += 1
c.used = stream_walk_no
}
prof_gen_add(c.count + 512)
}
if c == null {
missing = true
# Until the finer band is generated, show the coarser one this ground had
# a moment ago (same cell, next band out): approaching grass thins for a
# few frames instead of disappearing.
var b2 = band + 1
while c == null and b2 < 4 { c = stream_find(s, stream_key(cx, cz, b2)); b2 += 1 }
if c != null { c.used = stream_walk_no }
}
if c != null and c.count > 0 and l.count + c.count <= l.cap and stream_chunk_visible(s, cx, cz, c) {
let tg = gl_now_us()
mem_copy(mem_off(l.inst, l.count * INST_FLOATS * 4), c.data, c.count * INST_FLOATS * 4)
l.count += c.count
stream_us_gather = stream_us_gather + (gl_now_us() - tg)
}
}
}
cx += 1
}
cz += 1
}
ring += 1
}
s.pending = missing
stream_us_walk = stream_us_walk + (gl_now_us() - tw)
# force the layer to re-partition its (new) instances
l.view_gen = -1
}
# Only chunks that can be seen are gathered: a sphere around the chunk's footprint and
# height range, padded for the tallest cover and for casters just outside the frame
# whose short shadows still fall inside it.
function stream_chunk_visible(s: Stream, cx: int, cz: int, c: Chunk) -> bool {
let half = f_mul(s.size, F_HALF)
let wx = f_add(f_mul(fi(cx), s.size), half)
let wz = f_add(f_mul(fi(cz), s.size), half)
let hy = f_mul(f_sub(c.ymax, c.ymin), F_HALF)
let cy = f_add(c.ymin, hy)
let r = f_add(f_sqrt(f_add(f_mul(f_mul(half, half), F_TWO), f_mul(hy, hy))), fi(8))
return cam_sphere_visible(wx, cy, wz, r)
}
# what the caches hold, and whether they are being churned (R3D_PROF)
function stream_census() -> void {
if stream_all == null { return }
print("")
print(`ground-cover chunk caches (cap {string(STREAM_MAX_CHUNKS)} each, {string(stream_evictions)} evictions over the run):`)
var inst = 0
for i in 0 .. len(stream_all) {
let s = stream_all[i]
var n = 0
for k in 0 .. s.n { n += s.chunks[k].count }
inst += n
print(` stream kind {string(s.kind)}: {string(s.n)} chunks, {string(n)} instances`)
}
print(` {string(inst)} instances held, {string(inst * INST_FLOATS * 4 / 1024)} KB`)
}
function stream_update_all() -> void {
if stream_all == null { return }
stream_deadline = gl_now_us() + STREAM_BUDGET_US
for i in 0 .. len(stream_all) { stream_update(stream_all[i], cam_pos[0], cam_pos[2]) }
}

View file

@ -0,0 +1,706 @@
# ============================================================================
# terrain.ludic — the landscape: a height map generated on the GPU (R32F),
# read back for placement queries, drawn as a lifted grid with four scanned
# PBR materials blended by slope, altitude and the track mask.
# ============================================================================
var TERRAIN_HALF: int = 4096 # world half-size in metres (8 km square)
const TERRAIN_RES: int = 4096 # height-map texels per side (2 m over 8 km)
const TERRAIN_SHADOW_RES: int = 2048 # the baked height-field shadow / cloud mask
# CDLOD: the map is a quadtree of 32x32-cell patches; the leaf patch is 32 m (1 m cells)
const CD_G: int = 32 # cells per patch side
const CD_LEVELS: int = 9 # 32 m leaves .. 8192 m root
const CD_LEAVES: int = 256 # leaf patches per side (8192 / 32)
var cd_mesh: Mesh = null
var cd_range: words = null # float bits: how far each level is drawn
var cd_min: []words = null # per level: min height of each patch (float bits)
var cd_max: []words = null
var cd_draws: int = 0
var cd_far_draws: int = 0
var cd_near_draws: int = 0
var ter_force_far: bool = false
var ter_no_split: bool = false
var ter_force_near: bool = false
var ter_skip: bool = false
var ter_height_tex: int = 0
var ter_heights: words = null # CPU copy, float bits, TERRAIN_RES^2
var ter_reflect: bool = false # drawing the reflection: the mid mesh is plenty
var ter_prog: int = 0
# The far tier compiled on its own (FAR_ONLY). A patch that lies entirely beyond the
# near/far split is drawn with it: same pixels, a shader small enough to run wide.
var ter_prog_far: int = 0
var ter_prog_near: int = 0 # the detailed tier alone (NEAR_ONLY)
var ter_sun_prog: int = 0 # tersun.frag: sun visibility into a screen buffer
var ter_sun_tgt: Target = null # that buffer, at the frame's size
var ter_sun_refl: Target = null # and at the reflection's, which is smaller
var ter_sun_dumped: bool = false
var ter_sun_checked: bool = false
var ter_sun_done: bool = false # the caller already ran the pass (so it can time it)
var ter_sun_pass: bool = false # selection is drawing the visibility pass
var ter_sun_tex: int = 0
var ter_prog_cur: int = 0 # the program currently bound during selection
var ter_smooth: bool = false # generate the analytic test ground instead of a survey
var ter_far_split: int = 0 # metres: beyond this the terrain takes its cheap far path (R3D_TFAR)
var ter_far_band: int = 0 # half-width of the near/far blend (R3D_TBAND)
var ter_snow_line: int = 0
var ter_tex: words = null # 11 material textures (see terrain_bind)
var ter_ox: int = 0 # world x/z of the terrain centre (float bits)
var ter_oz: int = 0
var ter_dem_tex: int = 0 # a real height map (16-bit), or 0 for the procedural valley
var ter_dem_blur: int = 0 # gaussian texels applied to the survey (0 for lidar; ~3 for 30 m data)
var ter_dem_min: int = 0
var ter_dem_max: int = 0
var ter_dem_base: int = 0
var ter_ortho_tex: int = 0 # a photograph of the same window, draped with distance
var ter_carpet: int = 0 # the distant-grass carpet (carpet_bake), 0 = none
var ter_shadow_tex: int = 0 # height-field sun shadow: RG32F (lowest lit height, occluder distance)
var ter_shadow_yaw: int = 0x7fffffff # the sky yaw it was baked for
var ter_shadow_gen: int = -1 # the daylight generation it was baked for
var ter_shadow_prog: int = 0
function terrain_set_carpet(tex: int) -> void { ter_carpet = tex }
var ter_lake_level: int = 0 # a lake carved into the height map (float bits; ex = 0 → none)
var ter_lake_cx: int = 0
var ter_lake_cz: int = 0
var ter_lake_ex: int = 0
var ter_lake_ez: int = 0
# Carve a lake bed below `level` inside the ellipse (cx, cz) ± (ex, ez); call before r3d_init.
function terrain_lake(level: int, cx: int, cz: int, ex: int, ez: int) -> void {
ter_lake_level = level; ter_lake_cx = cx; ter_lake_cz = cz; ter_lake_ex = ex; ter_lake_ez = ez
}
# Use a real place: a 16-bit PNG height map plus its elevation range (metres). The
# elevation `base` becomes y = 0; `ox`/`oz` put the map's centre in the world.
var ter_ortho_px: pointer = null # the photograph on the CPU (RGB8, TERRAIN_RES^2) for placement rules
var ter_ortho_w: int = 0
var ter_ortho_c: int = 3
# the photograph's colour at world (x, z): packed 0xRRGGBB (0 outside the map)
function terrain_ortho(x: int, z: int) -> int {
if ter_ortho_px == null { return 0 }
if ter_o_scale == 0 { ter_o_scale = fr(ter_ortho_w, TERRAIN_HALF * 2) }
let scale = ter_o_scale
var ix = f_to_int(f_floor(f_mul(f_add(f_sub(x, ter_ox), fi(TERRAIN_HALF)), scale)))
var iz = f_to_int(f_floor(f_mul(f_add(f_sub(z, ter_oz), fi(TERRAIN_HALF)), scale)))
if ix < 0 { ix = 0 }; if iz < 0 { iz = 0 }
if ix > ter_ortho_w - 1 { ix = ter_ortho_w - 1 }; if iz > ter_ortho_w - 1 { iz = ter_ortho_w - 1 }
let o = (iz * ter_ortho_w + ix) * ter_ortho_c
return (ter_ortho_px[o] << 16) | (ter_ortho_px[o + 1] << 8) | ter_ortho_px[o + 2]
}
# The three classifiers below all read the same pixel. A caller that wants more than
# one should fetch the colour once with terrain_ortho() and use the *_of forms — the
# cover generator tests all three per candidate, so this is three fetches saved out of
# every four in the hottest loop in the program.
function ortho_green_of(c: int) -> int {
let r = (c >> 16) & 255; let g = (c >> 8) & 255; let b = c & 255
var v = g - max(r, b)
if v < 0 { v = 0 }
return f_min(fr(v, 22), F_ONE)
}
function ortho_scree_of(c: int) -> int {
if c == 0 { return F_ZERO }
let r = (c >> 16) & 255; let g = (c >> 8) & 255; let b = c & 255
let mx = max(r, max(g, b))
if g - max(r, b) > 2 or mx < 60 { return F_ZERO }
return F_ONE
}
function ortho_forest_of(c: int) -> int {
if c == 0 { return F_ZERO }
let r = (c >> 16) & 255; let g = (c >> 8) & 255; let b = c & 255
let mx = max(r, max(g, b))
if g - max(r, b) < 3 { return F_ZERO }
if mx <= 80 { return F_ONE }
if mx <= 105 { return F_HALF }
return F_ZERO
}
# how green the ground is in the photograph (0..1 float bits): meadow / forest vs rock, scree, water
function terrain_ortho_green(x: int, z: int) -> int {
let c = terrain_ortho(x, z)
let r = (c >> 16) & 255; let g = (c >> 8) & 255; let b = c & 255
var v = g - max(r, b) # green excess
if v < 0 { v = 0 }
return f_min(fr(v, 22), F_ONE)
}
# how grey and mid-bright (scree / pebbles / bare rock) the photograph is there (0..1)
# Bare ground only where most of a 50 m neighbourhood is bare: a single 10 m trail pixel
# must not place a boulder or bar a tree.
function terrain_ortho_scree(x: int, z: int) -> int {
var votes = 0
for j in 0 .. 5 { for i in 0 .. 5 { if ortho_scree_of(terrain_ortho(f_add(x, fi((i - 2) * 10)), f_add(z, fi((j - 2) * 10)))) != F_ZERO { votes += 1 } } }
if votes >= 15 { return F_ONE }
return F_ZERO
}
function terrain_ortho_scree_pixel(x: int, z: int) -> int {
let c = terrain_ortho(x, z)
if c == 0 { return F_ZERO }
let r = (c >> 16) & 255; let g = (c >> 8) & 255; let b = c & 255
let mx = max(r, max(g, b))
# bare ground: not green-dominant (grey scree, the maroon rock, the moraine's tan gravel), lit enough not to be water
if g - max(r, b) > 2 or mx < 60 { return F_ZERO }
return F_ONE
}
# dense conifer forest in the photograph: green-dominant and dark (the meadows are brighter)
function terrain_ortho_forest(x: int, z: int) -> int {
let c = terrain_ortho(x, z)
if c == 0 { return F_ZERO }
let r = (c >> 16) & 255; let g = (c >> 8) & 255; let b = c & 255
let mx = max(r, max(g, b))
if g - max(r, b) < 3 { return F_ZERO }
if mx <= 80 { return F_ONE }
if mx <= 105 { return F_HALF }
return F_ZERO
}
function terrain_use_ortho(path: string) -> void {
let px = png_decode(path)
if px == null { return }
ter_ortho_px = px
ter_ortho_w = tex_w
ter_ortho_c = tex_channels
ter_ortho_tex = tex_upload(px, true, true)
gl_bind_texture(GL_TEXTURE_2D, ter_ortho_tex)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_WRAP_S, GL_CLAMP_TO_EDGE)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_WRAP_T, GL_CLAMP_TO_EDGE)
}
function terrain_use_dem(path: string, emin: int, emax: int, base: int, ox: int, oz: int) -> void {
ter_dem_tex = tex_load(path, false)
ter_dem_min = emin; ter_dem_max = emax; ter_dem_base = base
ter_ox = ox; ter_oz = oz
gl_bind_texture(GL_TEXTURE_2D, ter_dem_tex)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_WRAP_S, GL_CLAMP_TO_EDGE)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_WRAP_T, GL_CLAMP_TO_EDGE)
}
function terrain_generate() -> void {
var ids0: words = null
var defs = ""
if ter_dem_tex != 0 { defs = "#define DEM\n" }
if ter_smooth { defs = "#define SMOOTH\n" }
let p = r3d_program("fullscreen.vert", "heightgen.frag", defs)
ter_height_tex = tex_target(TERRAIN_RES, TERRAIN_RES, GL_R32F, GL_RED, GL_FLOAT, GL_LINEAR)
let fbo = gl_framebuffer()
gl_bind_framebuffer(GL_FRAMEBUFFER, fbo)
gl_framebuffer_texture2d(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, ter_height_tex, 0)
gl_viewport(0, 0, TERRAIN_RES, TERRAIN_RES)
gl_disable(GL_DEPTH_TEST)
gl_use_program(p)
u_f(gl_uniform(p, "u_half"), fi(TERRAIN_HALF))
if ter_dem_tex != 0 {
r3d_bind_2d(p, "u_dem", 0, ter_dem_tex)
u_f(gl_uniform(p, "u_dem_min"), ter_dem_min)
u_f(gl_uniform(p, "u_dem_max"), ter_dem_max)
u_f(gl_uniform(p, "u_dem_base"), ter_dem_base)
u_f2(gl_uniform(p, "u_origin"), ter_ox, ter_oz)
u_f4(gl_uniform(p, "u_lake"), ter_lake_cx, ter_lake_cz, ter_lake_ex, ter_lake_ez)
u_f(gl_uniform(p, "u_lake_level"), ter_lake_level)
u_f(gl_uniform(p, "u_dem_blur"), ter_dem_blur)
}
mesh_draw(sky_fullscreen)
# second pass: R = height, GBA = the smooth surface normal, baked once (ternormal.frag)
let raw = ter_height_tex
ter_height_tex = tex_target(TERRAIN_RES, TERRAIN_RES, GL_RGBA32F, GL_RGBA, GL_FLOAT, GL_LINEAR)
gl_framebuffer_texture2d(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, ter_height_tex, 0)
let pn = r3d_program("fullscreen.vert", "ternormal.frag", "")
gl_use_program(pn)
r3d_bind_2d(pn, "u_src", 0, raw)
u_f(gl_uniform(pn, "u_half"), fi(TERRAIN_HALF))
mesh_draw(sky_fullscreen)
gl_delete_program(pn)
ids0 = gl_scratch(); ids0[0] = raw; gl_delete_textures(1, ids0)
# read the heights back for placement
ter_heights = words(TERRAIN_RES * TERRAIN_RES)
gl_bind_texture(GL_TEXTURE_2D, ter_height_tex)
gl_pixel_storei(GL_PACK_ALIGNMENT, 4)
gl_get_tex_image(GL_TEXTURE_2D, 0, GL_RED, GL_FLOAT, ter_heights)
gl_bind_framebuffer(GL_FRAMEBUFFER, 0)
let ids = gl_scratch()
ids[0] = fbo
gl_delete_framebuffers(1, ids)
gl_delete_program(p)
gl_check("terrain generate")
}
# The height field for shaders that place things on the ground (model.vert's u_ground)
function terrain_bind_height(p: int) -> void {
r3d_bind_2d(p, "u_ts_height", 5, ter_height_tex)
u_f(gl_uniform(p, "u_ts_half"), fi(TERRAIN_HALF))
u_f2(gl_uniform(p, "u_ts_origin"), ter_ox, ter_oz)
}
# Bake the height-field sun shadow (see tershadow.frag). Cheap enough to redo whenever
# the sun moves; r3d_frame calls it again when sky_set_yaw has changed the yaw.
function terrain_bake_shadow() -> void {
if sun_dir == null { return }
if ter_shadow_tex == 0 { ter_shadow_tex = tex_target(TERRAIN_SHADOW_RES, TERRAIN_SHADOW_RES, GL_RGBA32F, GL_RGBA, GL_FLOAT, GL_LINEAR) }
if ter_shadow_prog == 0 { ter_shadow_prog = r3d_program("fullscreen.vert", "tershadow.frag", "#define NOISE_ONLY\n") }
let p = ter_shadow_prog
let fbo = gl_framebuffer()
gl_bind_framebuffer(GL_FRAMEBUFFER, fbo)
gl_framebuffer_texture2d(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, ter_shadow_tex, 0)
gl_viewport(0, 0, TERRAIN_SHADOW_RES, TERRAIN_SHADOW_RES)
gl_disable(GL_DEPTH_TEST)
gl_disable(GL_BLEND)
gl_use_program(p)
r3d_bind_2d(p, "u_height", 0, ter_height_tex)
u_f(gl_uniform(p, "u_half"), fi(TERRAIN_HALF))
u_v3(gl_uniform(p, "u_sun"), sun_dir)
mesh_draw(sky_fullscreen)
gl_bind_framebuffer(GL_FRAMEBUFFER, 0)
let ids = gl_scratch()
ids[0] = fbo
gl_delete_framebuffers(1, ids)
ter_shadow_yaw = sky_yaw
gl_check("terrain shadow bake")
}
# height at world (x, z) — float bits, bilinear over the CPU copy
# The height and photograph scales are constants, but they were being recomputed —
# a fixed-point divide — on every call, and the cover generator calls terrain_height
# five times per candidate (once directly, four more inside slope_at) across hundreds
# of thousands of candidates per chunk. Hoisted, they cost nothing.
var ter_h_scale: int = 0
var ter_o_scale: int = 0
function terrain_height(x: int, z: int) -> int {
if ter_h_scale == 0 { ter_h_scale = fr(TERRAIN_RES, TERRAIN_HALF * 2) }
let scale = ter_h_scale
let fx = f_mul(f_add(f_sub(x, ter_ox), fi(TERRAIN_HALF)), scale)
let fz = f_mul(f_add(f_sub(z, ter_oz), fi(TERRAIN_HALF)), scale)
var ix = f_to_int(f_floor(fx)); var iz = f_to_int(f_floor(fz))
if ix < 0 { ix = 0 }; if iz < 0 { iz = 0 }
if ix > TERRAIN_RES - 2 { ix = TERRAIN_RES - 2 }; if iz > TERRAIN_RES - 2 { iz = TERRAIN_RES - 2 }
let tx = f_clamp(f_sub(fx, fi(ix)), F_ZERO, F_ONE)
let tz = f_clamp(f_sub(fz, fi(iz)), F_ZERO, F_ONE)
let h00 = ter_heights[iz * TERRAIN_RES + ix]
let h10 = ter_heights[iz * TERRAIN_RES + ix + 1]
let h01 = ter_heights[(iz + 1) * TERRAIN_RES + ix]
let h11 = ter_heights[(iz + 1) * TERRAIN_RES + ix + 1]
return f_lerp(f_lerp(h00, h10, tx), f_lerp(h01, h11, tx), tz)
}
# the same for Q16.16 callers
function terrain_height_fx(x: fixed, z: fixed) -> fixed { return f_fx(terrain_height(fl(x), fl(z))) }
# The height the terrain is DRAWN at: the cubic B-spline of the texels (heightSmooth in
# terrain.vert), not the bilinear read above. The two differ by up to half a metre on
# rough ground, which is the difference between a character standing on the meadow
# and one buried to the knee in it. Sixteen taps; for things that move, not for the
# thousands of placement queries a chunk makes.
var ter_bw: words = null
function terrain_height_smooth(x: int, z: int) -> int {
if ter_h_scale == 0 { ter_h_scale = fr(TERRAIN_RES, TERRAIN_HALF * 2) }
if ter_bw == null { ter_bw = words(8) }
let scale = ter_h_scale
let fx = f_sub(f_mul(f_add(f_sub(x, ter_ox), fi(TERRAIN_HALF)), scale), F_HALF)
let fz = f_sub(f_mul(f_add(f_sub(z, ter_oz), fi(TERRAIN_HALF)), scale), F_HALF)
let ix = f_to_int(f_floor(fx)); let iz = f_to_int(f_floor(fz))
let tx = f_clamp(f_sub(fx, fi(ix)), F_ZERO, F_ONE)
let tz = f_clamp(f_sub(fz, fi(iz)), F_ZERO, F_ONE)
# the four cubic B-spline weights of a fraction, over texels i-1 .. i+2
for a in 0 .. 2 {
var t = tx
if a == 1 { t = tz }
let t2 = f_mul(t, t); let t3 = f_mul(t2, t)
let one_t = f_sub(F_ONE, t)
let w0 = f_div(f_mul(f_mul(one_t, one_t), one_t), fi(6))
let w1 = f_div(f_add(f_sub(fi(4), f_mul(fi(6), t2)), f_mul(fi(3), t3)), fi(6))
let w3 = f_div(t3, fi(6))
let w2 = f_sub(f_sub(f_sub(F_ONE, w0), w1), w3)
ter_bw[a * 4] = w0; ter_bw[a * 4 + 1] = w1; ter_bw[a * 4 + 2] = w2; ter_bw[a * 4 + 3] = w3
}
var h = F_ZERO
for j in 0 .. 4 {
var rz = iz - 1 + j
if rz < 0 { rz = 0 }; if rz > TERRAIN_RES - 1 { rz = TERRAIN_RES - 1 }
var row = F_ZERO
for i in 0 .. 4 {
var rx = ix - 1 + i
if rx < 0 { rx = 0 }; if rx > TERRAIN_RES - 1 { rx = TERRAIN_RES - 1 }
row = f_add(row, f_mul(ter_bw[i], ter_heights[rz * TERRAIN_RES + rx]))
}
h = f_add(h, f_mul(ter_bw[4 + j], row))
}
return h
}
function terrain_load_textures() -> void {
ter_tex = words(15)
let a = r3d_assets + "/textures/"
ter_tex[0] = tex_load(a + "aerial_grass_rock_diff_2k.png", true)
ter_tex[1] = tex_load(a + "aerial_grass_rock_nor_gl_2k.png", false)
ter_tex[2] = tex_load(a + "aerial_grass_rock_arm_2k.png", false)
ter_tex[3] = tex_load(a + "grass_path_2_diff_2k.png", true)
ter_tex[4] = tex_load(a + "grass_path_2_nor_gl_2k.png", false)
ter_tex[5] = tex_load(a + "grass_path_2_arm_2k.png", false)
ter_tex[6] = tex_load(a + "gray_rocks_diff_2k.png", true)
ter_tex[7] = tex_load(a + "gray_rocks_nor_gl_2k.png", false)
ter_tex[8] = tex_load(a + "gray_rocks_arm_2k.png", false)
ter_tex[9] = tex_load(a + "snow_02_diff_2k.png", true)
ter_tex[10] = tex_load(a + "snow_02_nor_gl_2k.png", false)
ter_tex[11] = tex_load(a + "snow_02_arm_2k.png", false)
ter_tex[12] = tex_load(a + "aerial_grass_rock_disp_2k.png", false)
ter_tex[13] = tex_load(a + "cliff_side_diff_2k.png", true)
ter_tex[14] = tex_load(a + "cliff_side_nor_gl_2k.png", false)
if Os.has_env("R3D_TEXDBG") { for i in 0 .. 15 { print(`ter_tex[{string(i)}] = {string(ter_tex[i])}`) } }
}
function terrain_init() -> void {
terrain_generate()
terrain_bake_shadow()
terrain_load_textures()
cdlod_init()
ter_wire = Os.has_env("R3D_WIRE")
ter_force_far = Os.has_env("R3D_TFARONLY")
ter_no_split = Os.has_env("R3D_NOSPLIT")
ter_force_near = Os.has_env("R3D_TNEARONLY")
ter_skip = Os.has_env("R3D_NOTERRAIN")
var defs = ""
if r3d_debug_shadow { defs = "#define DEBUG_SHADOW\n" }
if Os.has_env("R3D_DEBUG_MAT") { defs = "#define DEBUG_MAT\n" }
if Os.has_env("R3D_DEBUG_NRM") { defs = "#define DEBUG_NRM\n" }
if Os.has_env("R3D_DEBUG_ALB") { defs = "#define DEBUG_ALB\n" }
# Elimination profiling. Measure these by FRAME TIME (prof_ft_report), not by the
# per-pass GPU timers: this driver's timer queries attribute a pass's fragment work
# almost arbitrarily, and will happily report a pass at a tenth of its cost.
# R3D_TFAST=1 the ground reads no sun visibility =2 every noise field at its mean
# =3 no scanned material taps =4 no survey photograph
# =5 the detailed tier at every distance =20 no photograph grain
# R3D_NOTERRAIN skips the ground entirely (what it costs); R3D_TNEARONLY / R3D_TFARONLY
# draw every patch with one tier's program (what each tier costs over a whole frame);
# R3D_NOSPLIT goes back to the single program that holds both tiers.
if Os.has_env("R3D_TFAST") { defs = defs + "#define TFAST_" + Os.env("R3D_TFAST") + "\n" }
ter_prog = r3d_program("terrain.vert", "terrain.frag", defs)
ter_prog_far = r3d_program("terrain.vert", "terrain.frag", defs + "#define FAR_ONLY\n")
ter_prog_near = r3d_program("terrain.vert", "terrain.frag", defs + "#define NEAR_ONLY\n")
ter_sun_prog = r3d_program("terrain.vert", "tersun.frag", "")
ter_far_split = fi(200)
if Os.has_env("R3D_TFAR") { ter_far_split = fi(Text.to_int(Os.env("R3D_TFAR"))) }
ter_far_band = fi(60)
if Os.has_env("R3D_TBAND") { ter_far_band = fi(Text.to_int(Os.env("R3D_TBAND"))) }
ter_snow_line = fi(880)
gl_check("terrain init")
}
# draw into the current cascade with the given light view-projection
# The terrain no longer casts into the shadow map: it shadows itself by marching its
# own height field in terrain.frag, which cannot produce the self-shadow grid a depth
# map does, and it saves drawing the whole grid five times a frame.
function terrain_draw_shadow(light_vp: words) -> void {
}
# Every per-frame uniform of one terrain program. Both tiers are bound up front so
# selection can switch between them per patch without re-binding anything but the node.
function terrain_bind_prog(p: int) -> void {
gl_use_program(p)
r3d_bind_2d(p, "u_height", 0, ter_height_tex)
r3d_bind_2d(p, "u_grass_d", 1, ter_tex[0]); r3d_bind_2d(p, "u_grass_n", 2, ter_tex[1]); r3d_bind_2d(p, "u_grass_a", 3, ter_tex[2])
# The cliff maps went with the dead cliff sample. Binding textures for uniforms the
# shader no longer declares leaves those units pointing at nothing, which the driver
# reports as an unloadable sampler and resolves as a zero texture.
# A scene with no photograph still has to bind something valid here: sampler unit
# pointed at texture 0 is an incomplete texture, which the driver reports as
# unloadable and which poisons sampling for the rest of the unit's stage.
var orthotex = ter_ortho_tex
if orthotex == 0 { orthotex = ter_tex[0] }
r3d_bind_2d(p, "u_ortho", 4, orthotex)
var oon = F_ZERO
if ter_ortho_tex != 0 { oon = F_ONE }
u_f(gl_uniform(p, "u_ortho_on"), oon)
r3d_bind_2d(p, "u_rock_d", 7, ter_tex[6]); r3d_bind_2d(p, "u_rock_n", 8, ter_tex[7]); r3d_bind_2d(p, "u_rock_a", 9, ter_tex[8])
r3d_bind_2d(p, "u_snow_d", 10, ter_tex[9])
if ter_carpet != 0 { r3d_bind_2d(p, "u_carpet", 11, ter_carpet); u_f(gl_uniform(p, "u_carpet_on"), F_ONE) }
else { r3d_bind_2d(p, "u_carpet", 11, ter_tex[0]); u_f(gl_uniform(p, "u_carpet_on"), F_ZERO) }
u_f(gl_uniform(p, "u_half"), fi(TERRAIN_HALF))
u_f(gl_uniform(p, "u_texel"), fr(1, TERRAIN_RES))
u_f(gl_uniform(p, "u_snow_line"), ter_snow_line)
var lake = fl(-100000.0)
if ter_lake_ex != 0 { lake = ter_lake_level }
u_f(gl_uniform(p, "u_lake_level"), lake)
u_mat4(gl_uniform(p, "u_view"), cam_view)
u_mat4(gl_uniform(p, "u_proj"), cam_proj)
u_f2(gl_uniform(p, "u_origin"), ter_ox, ter_oz)
u_f(gl_uniform(p, "u_far_split"), ter_far_split)
u_f(gl_uniform(p, "u_far_band"), ter_far_band)
sky_bind_lighting(p)
shadow_bind(p)
fog_bind(p)
u_v3(gl_uniform(p, "u_cam_pos"), cam_pos)
u_f(gl_uniform(p, "u_grid"), fi(CD_G))
# The ground reads its sun visibility out of the buffer tersun.frag filled, and has no
# use for the cascade array shadow_bind just put on this unit; leaving both bound under
# one unit is undefined ground, so the array comes off first.
gl_active_texture(GL_TEXTURE0 + 15)
gl_bind_texture(GL_TEXTURE_2D_ARRAY, 0)
r3d_bind_2d(p, "u_sunshadow", 15, ter_sun_tex)
}
# The visibility buffer for the size being drawn into. The reflection is rendered at its
# own (smaller) size, so it keeps its own.
#
# Neither owns a depth buffer. The pass BORROWS the depth the frame is about to be drawn
# with, so its rasterisation doubles as a depth prepass — and a borrowed texture must
# never be written into the Target, because a Target deletes whatever its `depth` names
# when it is freed. Storing it there deleted the frame's own depth buffer on the first
# resize (post_init had already made the replacement, and GL hands the freed name straight
# back, so the new one was deleted instead of the old). The scene framebuffer lost its
# depth attachment, the sky's fullscreen quad had nothing left to fail against, and it
# painted over the whole valley — with "gl error 1286" every frame from the passes whose
# attachment now named a deleted texture.
function terrain_sun_target(w: int, h: int) -> Target {
var t = ter_sun_tgt
if ter_reflect { t = ter_sun_refl }
if t == null or t.w != w or t.h != h {
target_free(t) # owns its colour, and nothing else
t = target_new(w, h, GL_R8, GL_RED, GL_UNSIGNED_BYTE, false, GL_NEAREST)
if ter_reflect { ter_sun_refl = t } else { ter_sun_tgt = t }
}
return t
}
# Rasterise the patches once with the small shader that reads the cascades (tersun.frag).
function terrain_sun_pass(w: int, h: int, depth: int) -> Target {
let t = terrain_sun_target(w, h)
# Attach the frame's depth afresh every pass. It is a different texture every time the
# screen-sized buffers are rebuilt — a resize, a fullscreen change — and an attachment
# naming a texture that has been deleted leaves this framebuffer incomplete, which is
# an error per draw and a pass that silently does nothing. One call a pass is cheaper
# than any scheme for noticing.
target_bind(t)
gl_framebuffer_texture2d(GL_FRAMEBUFFER, GL_DEPTH_ATTACHMENT, GL_TEXTURE_2D, depth, 0)
if not ter_sun_checked {
ter_sun_checked = true
let st = gl_check_framebuffer_status(GL_FRAMEBUFFER)
if st != GL_FRAMEBUFFER_COMPLETE { print(`r3d: sun-visibility framebuffer incomplete {st}`) }
}
gl_enable(GL_DEPTH_TEST)
gl_depth_func(GL_LESS)
gl_depth_mask(1)
# the depth is the frame's own and was cleared with it; only the visibility is cleared
gl_clear_color(1.0, 1.0, 1.0, 1.0) # unshadowed where nothing is drawn
gl_clear(GL_COLOR_BUFFER_BIT)
let p = ter_sun_prog
gl_use_program(p)
r3d_bind_2d(p, "u_height", 0, ter_height_tex)
u_f(gl_uniform(p, "u_half"), fi(TERRAIN_HALF))
u_mat4(gl_uniform(p, "u_view"), cam_view)
u_mat4(gl_uniform(p, "u_proj"), cam_proj)
u_f2(gl_uniform(p, "u_origin"), ter_ox, ter_oz)
u_v3(gl_uniform(p, "u_cam_pos"), cam_pos)
u_f(gl_uniform(p, "u_grid"), fi(CD_G))
u_f(gl_uniform(p, "u_far_split"), ter_far_split)
u_f(gl_uniform(p, "u_far_band"), ter_far_band)
u_f(gl_uniform(p, "u_clip_y"), r3d_clip_y)
shadow_bind(p)
sky_bind_lighting(p)
ter_sun_pass = true
gl_bind_vertex_array(cd_mesh.vao)
cdlod_select(CD_LEVELS - 1, 0, 0)
ter_sun_pass = false
if Os.has_env("R3D_DUMP_SUN") and not ter_sun_dumped and not ter_reflect { ter_sun_dumped = true; tex_dump(t.color, w, h, "build/dbg_sun.ppm") }
return t
}
function terrain_sun_prepare() -> void {
ter_sun_tex = terrain_sun_pass(post_w, post_h, post_hdr.depth).color
ter_sun_done = true
}
function terrain_draw() -> void {
if ter_skip { return }
# The sun visibility first, into its own buffer; the shading pass looks it up per pixel.
# The pass binds its own framebuffer, so the caller's target is restored afterwards —
# the reflection's, or the scene's, without disturbing what is already drawn in it.
var vw = post_w
var vh = post_h
if ter_reflect { vw = water_refl.w; vh = water_refl.h }
if not ter_sun_done {
var dep = post_hdr.depth
if ter_reflect { dep = water_refl.depth }
ter_sun_tex = terrain_sun_pass(vw, vh, dep).color
}
ter_sun_done = false
if ter_reflect { target_bind(water_refl) }
else {
target_bind(post_hdr)
if post_ms_fbo != 0 { gl_bind_framebuffer(GL_FRAMEBUFFER, post_ms_fbo) }
}
gl_enable(GL_DEPTH_TEST)
# the prepass already laid this geometry's depth down: only the frontmost fragment of
# each pixel has anything to shade, and it meets that depth exactly
gl_depth_func(GL_LEQUAL)
gl_depth_mask(1)
terrain_bind_prog(ter_prog)
terrain_bind_prog(ter_prog_far)
terrain_bind_prog(ter_prog_near)
ter_prog_cur = 0
cd_draws = 0
cd_far_draws = 0
cd_near_draws = 0
gl_bind_vertex_array(cd_mesh.vao)
if ter_wire { gl_polygon_mode(GL_FRONT_AND_BACK, GL_LINE) }
cdlod_select(CD_LEVELS - 1, 0, 0)
if ter_wire { gl_polygon_mode(GL_FRONT_AND_BACK, GL_FILL) }
gl_depth_func(GL_LESS)
if r3d_debug and not ter_printed { ter_printed = true; print(`cdlod patches drawn: {cd_draws} (far {cd_far_draws}, near {cd_near_draws}, band {cd_draws - cd_far_draws - cd_near_draws})`) }
}
var ter_wire: bool = false
var ter_printed: bool = false
# ---- CDLOD --------------------------------------------------------------------------
# One 32x32 patch mesh (a_xz in 0..1) drawn once per selected quadtree node; the vertex
# shader places, scales and morphs it. Levels are drawn out to cd_range[k] = 48 * 2^k m,
# so cells are 1 m within 48 m, 2 m to 96 m, 4 m to 192 m ... 256 m at the root.
function cdlod_init() -> void {
let m = new Mesh
m.vao = gl_vao()
let n = CD_G + 1
let v = gl_floats(n * n * 2)
var k = 0
for j in 0 .. n { for i in 0 .. n { gl_put_bits(v, k, fr(i, CD_G)); gl_put_bits(v, k + 1, fr(j, CD_G)); k += 2 } }
m.vbo = gl_buffer()
gl_bind_buffer(GL_ARRAY_BUFFER, m.vbo)
gl_buffer_data(GL_ARRAY_BUFFER, gl_bytes_of(n * n * 2), v, GL_STATIC_DRAW)
gl_enable_vertex_attrib_array(0); gl_vertex_attrib_pointer(0, 2, GL_FLOAT, 0, 8, null)
free(v)
let ni = CD_G * CD_G * 6
let idx = words(ni)
k = 0
for j in 0 .. CD_G {
for i in 0 .. CD_G {
let a = j * n + i
idx[k] = a; idx[k + 1] = a + n; idx[k + 2] = a + 1
idx[k + 3] = a + 1; idx[k + 4] = a + n; idx[k + 5] = a + n + 1
k += 6
}
}
m.ebo = gl_buffer()
gl_bind_buffer(GL_ELEMENT_ARRAY_BUFFER, m.ebo)
gl_buffer_data(GL_ELEMENT_ARRAY_BUFFER, ni * 4, idx, GL_STATIC_DRAW)
free(idx)
m.count = ni
gl_bind_vertex_array(0)
cd_mesh = m
cd_range = words(CD_LEVELS)
var r = fi(48)
if Os.has_env("R3D_CD_R0") { r = fi(Text.to_int(Os.env("R3D_CD_R0"))) }
for l in 0 .. CD_LEVELS { cd_range[l] = r; r = f_mul(r, F_TWO) }
cdlod_bounds()
}
# min/max height per patch at every level, from the CPU copy of the height field
function cdlod_bounds() -> void {
cd_min = new []words; cd_max = new []words
let t = TERRAIN_RES / CD_LEAVES # texels per leaf patch side
var n = CD_LEAVES
var lo = words(n * n); var hi = words(n * n)
for j in 0 .. n {
for i in 0 .. n {
var mn = fi(100000); var mx = fi(-100000)
for y in 0 .. t + 1 {
let ty = min(j * t + y, TERRAIN_RES - 1)
for x in 0 .. t + 1 {
let tx = min(i * t + x, TERRAIN_RES - 1)
let h = ter_heights[ty * TERRAIN_RES + tx]
mn = f_min(mn, h); mx = f_max(mx, h)
}
}
lo[j * n + i] = mn; hi[j * n + i] = mx
}
}
push(cd_min, lo); push(cd_max, hi)
while n > 1 {
let m = n / 2
let plo = words(m * m); let phi = words(m * m)
for j in 0 .. m {
for i in 0 .. m {
let a = (2 * j) * n + 2 * i
plo[j * m + i] = f_min(f_min(lo[a], lo[a + 1]), f_min(lo[a + n], lo[a + n + 1]))
phi[j * m + i] = f_max(f_max(hi[a], hi[a + 1]), f_max(hi[a + n], hi[a + n + 1]))
}
}
push(cd_min, plo); push(cd_max, phi)
lo = plo; hi = phi; n = m
}
}
# does the patch's box come within r of the camera?
function cd_within(x0: int, z0: int, size: int, ymin: int, ymax: int, r: int) -> bool {
let dx = f_max(f_max(f_sub(x0, cam_pos[0]), f_sub(cam_pos[0], f_add(x0, size))), F_ZERO)
let dz = f_max(f_max(f_sub(z0, cam_pos[2]), f_sub(cam_pos[2], f_add(z0, size))), F_ZERO)
let dy = f_max(f_max(f_sub(ymin, cam_pos[1]), f_sub(cam_pos[1], ymax)), F_ZERO)
return f_ls(f_add(f_add(f_mul(dx, dx), f_mul(dz, dz)), f_mul(dy, dy)), f_mul(r, r))
}
# is the patch's box entirely inside r of the camera? (its farthest corner is within r)
function cd_inside(x0: int, z0: int, size: int, ymin: int, ymax: int, r: int) -> bool {
let x1 = f_add(x0, size)
let z1 = f_add(z0, size)
let dx = f_max(f_abs(f_sub(cam_pos[0], x0)), f_abs(f_sub(cam_pos[0], x1)))
let dz = f_max(f_abs(f_sub(cam_pos[2], z0)), f_abs(f_sub(cam_pos[2], z1)))
let dy = f_max(f_abs(f_sub(cam_pos[1], ymin)), f_abs(f_sub(cam_pos[1], ymax)))
return f_ls(f_add(f_add(f_mul(dx, dx), f_mul(dz, dz)), f_mul(dy, dy)), f_mul(r, r))
}
function cdlod_draw(level: int, ix: int, iz: int) -> void {
let size = fi(32 << level)
let x0 = f_add(f_sub(ter_ox, fi(TERRAIN_HALF)), f_mul(fi(ix), size))
let z0 = f_add(f_sub(ter_oz, fi(TERRAIN_HALF)), f_mul(fi(iz), size))
# Which tier can run inside this patch. A patch that never comes within the split takes
# the cheap tier at every pixel; one that lies wholly inside it takes the detailed tier
# at every pixel. Only a patch that straddles the band needs the program that holds both
# and cross-fades between them — and there are few of those, one ring of them.
if ter_sun_pass {
let t = gl_scratch()
t[0] = x0; t[1] = z0; t[2] = size
gl_uniform3fv(gl_uniform(ter_sun_prog, "u_node"), 1, t)
var st0 = F_ZERO
if level > 0 { st0 = cd_range[level - 1] }
u_f2(gl_uniform(ter_sun_prog, "u_morph"), f_lerp(st0, cd_range[level], fl(0.7)), cd_range[level])
gl_draw_elements(GL_TRIANGLES, cd_mesh.count, GL_UNSIGNED_INT, null)
return
}
let n = CD_LEAVES >> level
let ymin = cd_min[level][iz * n + ix]
let ymax = cd_max[level][iz * n + ix]
var p = ter_prog
if not cd_within(x0, z0, size, ymin, ymax, f_add(ter_far_split, ter_far_band)) { p = ter_prog_far }
else if cd_inside(x0, z0, size, ymin, ymax, f_sub(ter_far_split, ter_far_band)) { p = ter_prog_near }
if ter_force_far { p = ter_prog_far }
if ter_no_split { p = ter_prog }
if ter_force_near { p = ter_prog_near }
if p == ter_prog_far { cd_far_draws += 1 }
if p == ter_prog_near { cd_near_draws += 1 }
if p != ter_prog_cur { gl_use_program(p); ter_prog_cur = p }
let t = gl_scratch()
t[0] = x0; t[1] = z0; t[2] = size
gl_uniform3fv(gl_uniform(p, "u_node"), 1, t)
var start = F_ZERO
if level > 0 { start = cd_range[level - 1] }
start = f_lerp(start, cd_range[level], fl(0.7))
u_f2(gl_uniform(p, "u_morph"), start, cd_range[level])
gl_draw_elements(GL_TRIANGLES, cd_mesh.count, GL_UNSIGNED_INT, null)
cd_draws += 1
}
# Strugar's selection: a node is drawn at its own level unless it is close enough to need
# its children, in which case each child either selects itself or is drawn at this level.
function cdlod_select(level: int, ix: int, iz: int) -> bool {
let n = CD_LEAVES >> level
let size = fi(32 << level)
let x0 = f_add(f_sub(ter_ox, fi(TERRAIN_HALF)), f_mul(fi(ix), size))
let z0 = f_add(f_sub(ter_oz, fi(TERRAIN_HALF)), f_mul(fi(iz), size))
let ymin = cd_min[level][iz * n + ix]
let ymax = cd_max[level][iz * n + ix]
if not cd_within(x0, z0, size, ymin, ymax, cd_range[level]) { return false }
let half = f_mul(size, F_HALF)
let cy = f_mul(f_add(ymin, ymax), F_HALF)
let rad = f_sqrt(f_add(f_mul(f_mul(half, half), F_TWO), f_mul(f_mul(f_sub(ymax, cy), f_sub(ymax, cy)), F_ONE)))
if not cam_sphere_visible(f_add(x0, half), cy, f_add(z0, half), f_add(rad, fi(2))) { return true }
if level == 0 { cdlod_draw(0, ix, iz); return true }
if not cd_within(x0, z0, size, ymin, ymax, cd_range[level - 1]) { cdlod_draw(level, ix, iz); return true }
for c in 0 .. 4 {
let cx = ix * 2 + (c & 1); let cz = iz * 2 + (c >> 1)
if not cdlod_select(level - 1, cx, cz) { cdlod_draw(level - 1, cx, cz) }
}
return true
}

View file

@ -0,0 +1,472 @@
# ============================================================================
# texture.ludic — images for the GPU: PNG (8- and 16-bit, any colour type) and
# Radiance .hdr (RGBE) decoding straight into OpenGL textures.
#
# The engine's own PNG reader (image.ludic) expands to 8-bit 0xAARRGGBB for the
# 2D framebuffer; a renderer wants the file's real sample depth — normal and
# displacement maps ship as 16-bit — so this decoder keeps 16-bit samples and
# uploads them as GL_UNSIGNED_SHORT (big-endian, with GL_UNPACK_SWAP_BYTES) into
# RGB16 / R16 textures, and 8-bit ones into sRGB8 or RGB8 as the caller says.
# ============================================================================
const GL_TEXTURE_MAX_ANISOTROPY_EXT: int = 0x84FE
var tex_w: int = 0 # the last decoded image
var tex_h: int = 0
var tex_channels: int = 0
var tex_depth: int = 0 # bits per sample (8 or 16)
var tex_file_len: int = 0
var tex_anisotropy: fixed = 16.0
function r3d_read_file(path: pointer) -> pointer {
let f = file_open(path, "rb")
if f == null { return null }
file_seek(f, 0, 2)
let n = file_tell(f)
file_seek(f, 0, 0)
if n <= 0 { file_close(f); return null }
let buf = bytes(n + 8)
file_read(f, buf, n)
file_close(f)
tex_file_len = n
return buf
}
function be32(b: pointer, at: int) -> int {
return (b[at] << 24) | (b[at + 1] << 16) | (b[at + 2] << 8) | b[at + 3]
}
function tag4(b: pointer, at: int, a: int, c: int, d: int, e: int) -> bool {
return b[at] == a and b[at + 1] == c and b[at + 2] == d and b[at + 3] == e
}
# Decode a PNG into tightly packed scanlines of raw samples (PNG byte order:
# 16-bit samples big-endian). Sets tex_w / tex_h / tex_channels / tex_depth.
# Indexed and sub-byte greyscale files are expanded to 8-bit RGB / grey.
# Reverse one scanline's PNG filter in place (spec 9.2). The filter type is
# loop-invariant, so it is resolved once here rather than per byte, and the
# leading `fbpp` bytes (where the left neighbour is zero by definition) run as
# their own prologue instead of costing a bounds test on every byte of the image.
# The caller keeps a zeroed scanline in front of row 0, so `prev` is always a real
# row and every filter has exactly one code path — no first-row special cases to
# get wrong or to leave untested.
function png_unfilter(raw: pointer, cur: int, prev: int, stride: int, fbpp: int, ft: int) -> void {
if ft == 0 { return }
var first = fbpp
if first > stride { first = stride }
var x = 0
if ft == 1 {
x = fbpp
while x < stride { raw[cur + x] = ((raw[cur + x] + raw[cur + x - fbpp]) & 255); x += 1 }
return
}
if ft == 2 {
x = 0
while x < stride { raw[cur + x] = ((raw[cur + x] + raw[prev + x]) & 255); x += 1 }
return
}
if ft == 3 {
x = 0
while x < first { raw[cur + x] = ((raw[cur + x] + raw[prev + x] / 2) & 255); x += 1 }
while x < stride { raw[cur + x] = ((raw[cur + x] + (raw[cur + x - fbpp] + raw[prev + x]) / 2) & 255); x += 1 }
return
}
if ft == 4 {
x = 0
while x < first { raw[cur + x] = ((raw[cur + x] + raw[prev + x]) & 255); x += 1 }
while x < stride {
let a = raw[cur + x - fbpp]
let b = raw[prev + x]
let c = raw[prev + x - fbpp]
let p = a + b - c
let pa = abs(p - a)
let pb = abs(p - b)
let pc = abs(p - c)
var pick = c
if pb <= pc { pick = b }
if pa <= pb and pa <= pc { pick = a }
raw[cur + x] = ((raw[cur + x] + pick) & 255)
x += 1
}
}
}
function png_decode(path: pointer) -> pointer {
let d = r3d_read_file(path)
if d == null { print(`png: cannot read {path}`); return null }
let size = tex_file_len
if size < 8 or d[0] != 137 or d[1] != 80 { free(d); print(`png: not a png: {path}`); return null }
var w = 0; var h = 0; var bd = 0; var ct = 0
let plte = bytes(768)
let idat = bytes(size)
var idlen = 0
var i = 8
var done = false
while not done {
if i + 8 > size { done = true; continue }
let ln = be32(d, i)
let typ = i + 4
let body = i + 8
if ln < 0 or body + ln > size { done = true; continue }
if tag4(d, typ, 73, 72, 68, 82) { w = be32(d, body); h = be32(d, body + 4); bd = d[body + 8]; ct = d[body + 9] }
if tag4(d, typ, 80, 76, 84, 69) { let m = min(ln, 768); for k in 0 .. m { plte[k] = d[body + k] } }
if tag4(d, typ, 73, 68, 65, 84) { mem_copy(mem_off(idat, idlen), mem_off(d, body), ln); idlen += ln }
if tag4(d, typ, 73, 69, 78, 68) { done = true }
i = i + 12 + ln
}
if w <= 0 or h <= 0 { free(d); free(idat); free(plte); return null }
var channels = 1
if ct == 2 { channels = 3 }
if ct == 4 { channels = 2 }
if ct == 6 { channels = 4 }
let bppbits = bd * channels
var fbpp = (bppbits + 7) / 8
if fbpp < 1 { fbpp = 1 }
let stride = (w * bppbits + 7) / 8
let rawlen = h * (stride + 1)
# one zeroed scanline in front of the data, so row 0's "row above" is real
let raw = bytes(stride + rawlen + 8)
for z in 0 .. stride { raw[z] = 0 }
if z_uncompress(idat, idlen, mem_off(raw, stride), rawlen) < 0 { free(d); free(idat); free(raw); free(plte); print(`png: inflate failed: {path}`); return null }
free(d); free(idat)
# reverse the per-scanline filters in place, then pack rows without the filter byte
var y = 0
while y < h {
let line = stride + y * (stride + 1)
png_unfilter(raw, line + 1, line + 1 - (stride + 1), stride, fbpp, raw[line])
y += 1
}
var out: pointer = null
if (ct == 3) or (bd < 8) {
# expand palette / sub-byte grey to 8-bit RGB (palette) or 8-bit grey
let maxv = (1 << bd) - 1
var oc = 1
if ct == 3 { oc = 3 }
out = bytes(w * h * oc)
for yy in 0 .. h {
let row = stride + yy * (stride + 1) + 1
for x in 0 .. w {
let bp = x * bd
let idx = ((raw[row + bp / 8] >> (8 - bd - bp % 8)) & maxv)
if ct == 3 { out[(yy * w + x) * 3] = plte[idx * 3]; out[(yy * w + x) * 3 + 1] = plte[idx * 3 + 1]; out[(yy * w + x) * 3 + 2] = plte[idx * 3 + 2] }
else { out[yy * w + x] = idx * 255 / maxv }
}
}
channels = oc
bd = 8
free(raw)
} else {
out = bytes(h * stride + 8)
for yy in 0 .. h { mem_copy(mem_off(out, yy * stride), mem_off(raw, stride + yy * (stride + 1) + 1), stride) }
free(raw)
}
free(plte)
tex_w = w; tex_h = h; tex_channels = channels; tex_depth = bd
return out
}
# Edge padding for cut-out atlases: pixels darker than `thresh` (the unused
# background) take the mean of their lit neighbours, repeated `passes` times, so
# mipmaps and bilinear taps never pull black into the blades. 8-bit RGB/RGBA only.
function tex_dilate(px: pointer, thresh: int, passes: int) -> void {
if tex_depth != 8 or tex_channels < 3 { return }
let w = tex_w; let h = tex_h; let c = tex_channels
let mask = bytes(w * h)
var i = 0
while i < w * h { let o = i * c; if px[o] + px[o + 1] + px[o + 2] < thresh { mask[i] = 1 } else { mask[i] = 0 }; i += 1 }
let next = bytes(w * h)
for pass in 0 .. passes {
mem_copy(next, mask, w * h)
var y = 0
while y < h {
var x = 0
while x < w {
let k = y * w + x
if mask[k] == 1 {
var r = 0; var g = 0; var b = 0; var n = 0
if x > 0 and mask[k - 1] == 0 { let o = (k - 1) * c; r += px[o]; g += px[o + 1]; b += px[o + 2]; n += 1 }
if x < w - 1 and mask[k + 1] == 0 { let o = (k + 1) * c; r += px[o]; g += px[o + 1]; b += px[o + 2]; n += 1 }
if y > 0 and mask[k - w] == 0 { let o = (k - w) * c; r += px[o]; g += px[o + 1]; b += px[o + 2]; n += 1 }
if y < h - 1 and mask[k + w] == 0 { let o = (k + w) * c; r += px[o]; g += px[o + 1]; b += px[o + 2]; n += 1 }
if n > 0 { let o = k * c; px[o] = r / n; px[o + 1] = g / n; px[o + 2] = b / n; next[k] = 0 }
}
x += 1
}
y += 1
}
mem_copy(mask, next, w * h)
}
free(mask); free(next)
}
# Upload the last-decoded samples as a 2D texture. srgb: colour data (8-bit only).
function tex_upload(px: pointer, srgb: bool, mips: bool) -> int {
let id = gl_texture()
gl_bind_texture(GL_TEXTURE_2D, id)
var fmt = GL_RED
if tex_channels == 2 { fmt = GL_RG }
if tex_channels == 3 { fmt = GL_RGB }
if tex_channels == 4 { fmt = GL_RGBA }
var ifmt = GL_R8
var ty = GL_UNSIGNED_BYTE
if tex_depth == 16 {
ty = GL_UNSIGNED_SHORT
ifmt = GL_R16
if tex_channels == 2 { ifmt = GL_RG16 }
if tex_channels == 3 { ifmt = GL_RGB16 }
if tex_channels == 4 { ifmt = GL_RGBA16 }
gl_pixel_storei(GL_UNPACK_SWAP_BYTES, 1)
} else {
if tex_channels == 2 { ifmt = GL_RG8 }
if tex_channels == 3 { ifmt = GL_RGB8; if srgb { ifmt = GL_SRGB8 } }
if tex_channels == 4 { ifmt = GL_RGBA8; if srgb { ifmt = GL_SRGB8_ALPHA8 } }
gl_pixel_storei(GL_UNPACK_SWAP_BYTES, 0)
}
gl_pixel_storei(GL_UNPACK_ALIGNMENT, 1)
gl_tex_image2d(GL_TEXTURE_2D, 0, ifmt, tex_w, tex_h, 0, fmt, ty, px)
gl_pixel_storei(GL_UNPACK_SWAP_BYTES, 0)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_WRAP_S, GL_REPEAT)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_WRAP_T, GL_REPEAT)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_MAG_FILTER, GL_LINEAR)
if mips {
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_MIN_FILTER, GL_LINEAR_MIPMAP_LINEAR)
gl_generate_mipmap(GL_TEXTURE_2D)
gl_tex_parameterf(GL_TEXTURE_2D, GL_TEXTURE_MAX_ANISOTROPY_EXT, tex_anisotropy)
} else {
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_MIN_FILTER, GL_LINEAR)
}
return id
}
# Load a PNG as a mipmapped, anisotropic texture (0 on failure). srgb for albedo.
function tex_load(path: pointer, srgb: bool) -> int { return tex_load_ex(path, srgb, 0) }
# ... with `dilate` passes of edge padding for a cut-out atlas (0 = none)
function tex_load_ex(path: pointer, srgb: bool, dilate: int) -> int {
let px = png_decode(path)
if px == null { return 0 }
if dilate > 0 { tex_dilate(px, 60, dilate) }
let id = tex_upload(px, srgb, true)
free(px)
return id
}
# A small solid-colour fallback texture (linear rgb 0..255), for missing maps.
function tex_solid(r: int, g: int, b: int, a: int) -> int {
let px = bytes(16)
for i in 0 .. 4 { px[i * 4] = r; px[i * 4 + 1] = g; px[i * 4 + 2] = b; px[i * 4 + 3] = a }
tex_w = 2; tex_h = 2; tex_channels = 4; tex_depth = 8
let id = tex_upload(px, false, false)
free(px)
return id
}
# ---- Radiance .hdr (RGBE, new-style RLE) -> RGB float bits -------------------------
var hdr_max_lum: int = 0 # float bits of the brightest texel (sun finding)
var hdr_max_x: int = 0
var hdr_max_y: int = 0
var hdr_sun_r: int = 0 # irradiance (float bits) of everything above the IBL clip: the sun
var hdr_sun_g: int = 0
var hdr_sun_b: int = 0
var hdr_clip: int = 0 # float bits; texels above this (per channel) feed the sun, not the IBL
function hdr_decode(path: pointer) -> words {
let d = r3d_read_file(path)
if d == null { print(`hdr: cannot read {path}`); return null }
let size = tex_file_len
# header: lines until an empty line, then "-Y h +X w"
var i = 0
var blank = false
while i < size and not blank {
if d[i] == 10 and d[i + 1] == 10 { blank = true; i += 2 }
else { i += 1 }
}
# parse "-Y <h> +X <w>"
var h = 0; var w = 0
i += 3
while d[i] >= '0' and d[i] <= '9' { h = h * 10 + (d[i] - 48); i += 1 }
i += 4
while d[i] >= '0' and d[i] <= '9' { w = w * 10 + (d[i] - 48); i += 1 }
i += 1
if w <= 0 or h <= 0 { free(d); print(`hdr: bad header {path}`); return null }
let out = words(w * h * 3)
let line = bytes(w * 4)
var maxl = 0
if hdr_clip == 0 { hdr_clip = fi(20) }
var sr = F_ZERO; var sg = F_ZERO; var sb = F_ZERO
var skye = F_ZERO # sky irradiance on an upward face (clipped part only)
let dphi = f_div(f_mul(F_TWO, F_PI), fi(w))
let dth = f_div(F_PI, fi(h))
var y = 0
while y < h {
if d[i] == 2 and d[i + 1] == 2 and (d[i + 2] & 128) == 0 {
i += 4
for c in 0 .. 4 {
var x = 0
while x < w {
var n = d[i]; i += 1
if n > 128 {
n -= 128
let v = d[i]; i += 1
for k in 0 .. n { line[(x + k) * 4 + c] = v }
x += n
} else {
for k in 0 .. n { line[(x + k) * 4 + c] = d[i + k] }
i += n
x += n
}
}
}
} else {
for x in 0 .. w { for c in 0 .. 4 { line[x * 4 + c] = d[i + x * 4 + c] } }
i += w * 4
}
let sinth = f_sin(f_mul(f_add(fi(y), F_HALF), dth))
let domega = f_mul(f_mul(dphi, dth), sinth)
for x in 0 .. w {
let e = line[x * 4 + 3]
let o = (y * w + x) * 3
if e == 0 { out[o] = 0; out[o + 1] = 0; out[o + 2] = 0 }
else {
let sh = e - 136
let vr = f_ldexp(f_from_int(line[x * 4]), sh)
let vg = f_ldexp(f_from_int(line[x * 4 + 1]), sh)
let vb = f_ldexp(f_from_int(line[x * 4 + 2]), sh)
# the texture is capped at what a half-float holds; the sun is integrated uncapped
out[o] = f_min(vr, fi(60000))
out[o + 1] = f_min(vg, fi(60000))
out[o + 2] = f_min(vb, fi(60000))
let lum = f_add(f_add(vr, vg), vb)
if f_lt(maxl, lum) != 0 { maxl = lum; hdr_max_x = x; hdr_max_y = y }
if y < h / 2 { skye = f_add(skye, f_mul(f_mul(f_min(vg, hdr_clip), f_cos(f_mul(f_add(fi(y), F_HALF), dth))), domega)) }
if f_lt(hdr_clip, vg) != 0 or f_lt(hdr_clip, vr) != 0 {
sr = f_add(sr, f_mul(f_max(f_sub(vr, hdr_clip), F_ZERO), domega))
sg = f_add(sg, f_mul(f_max(f_sub(vg, hdr_clip), F_ZERO), domega))
sb = f_add(sb, f_mul(f_max(f_sub(vb, hdr_clip), F_ZERO), domega))
}
}
}
y += 1
}
free(line); free(d)
tex_w = w; tex_h = h; tex_channels = 3; tex_depth = 32
hdr_max_lum = maxl
hdr_sun_r = sr; hdr_sun_g = sg; hdr_sun_b = sb
print(`hdr: peak/1000 {f_fx(f_div(maxl, fi(1000)))} sky irradiance(up) {f_fx(skye)} sun irradiance {f_fx(sg)} (Q16.16 = /65536)`)
return out
}
# Load an equirectangular .hdr as an RGB16F texture with mips (clamped in v).
function tex_load_hdr(path: pointer) -> int {
let px = hdr_decode(path)
if px == null { return 0 }
let id = gl_texture()
gl_bind_texture(GL_TEXTURE_2D, id)
gl_pixel_storei(GL_UNPACK_ALIGNMENT, 4)
gl_tex_image2d(GL_TEXTURE_2D, 0, GL_RGB16F, tex_w, tex_h, 0, GL_RGB, GL_FLOAT, px)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_WRAP_S, GL_REPEAT)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_WRAP_T, GL_CLAMP_TO_EDGE)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_MAG_FILTER, GL_LINEAR)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_MIN_FILTER, GL_LINEAR_MIPMAP_LINEAR)
gl_generate_mipmap(GL_TEXTURE_2D)
free(px)
return id
}
# An empty render-target texture of the given internal format (no mips, clamped).
function tex_target(w: int, h: int, ifmt: int, fmt: int, ty: int, filter: int) -> int {
let id = gl_texture()
gl_bind_texture(GL_TEXTURE_2D, id)
gl_tex_image2d(GL_TEXTURE_2D, 0, ifmt, w, h, 0, fmt, ty, null)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_WRAP_S, GL_CLAMP_TO_EDGE)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_WRAP_T, GL_CLAMP_TO_EDGE)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_MAG_FILTER, filter)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_MIN_FILTER, filter)
return id
}
var tex_dump_alpha: bool = false
# Debug: the brightest texel of an RGBA float texture and where it is.
function tex_max(tex: int, w: int, h: int, tag: pointer) -> void {
let buf = words(w * h * 4)
gl_bind_texture(GL_TEXTURE_2D, tex)
gl_pixel_storei(GL_PACK_ALIGNMENT, 4)
gl_get_tex_image(GL_TEXTURE_2D, 0, GL_RGBA, GL_FLOAT, buf)
var best = F_ZERO; var bx = 0; var by = 0
var i = 0
while i < w * h {
let v = f_max(buf[i * 4], f_max(buf[i * 4 + 1], buf[i * 4 + 2]))
if f_gt(v, best) { best = v; bx = i % w; by = i / w }
i += 1
}
print(`{tag}: max {f_fx(f_div(best, fi(100)))}/100 at {bx} {h - 1 - by} (top-down)`)
let o = (by * w + bx) * 4
let big = fi(65000)
let finite = f_ls(best, big)
print(` rgba (clamped/100): {f_fx(f_div(f_min(buf[o], big), fi(100)))} {f_fx(f_div(f_min(buf[o + 1], big), fi(100)))} {f_fx(f_div(f_min(buf[o + 2], big), fi(100)))} {f_fx(f_min(buf[o + 3], big))} finite {finite} bits {buf[o]}`)
free(buf)
}
# Debug: write a 2D texture's level 0 (RGBA8, alpha dropped) as a binary PPM.
function tex_dump(tex: int, w: int, h: int, path: pointer) -> void {
let f = file_open(path, "wb")
if f == null { return }
let buf = bytes(w * h * 4)
gl_bind_texture(GL_TEXTURE_2D, tex)
gl_pixel_storei(GL_PACK_ALIGNMENT, 1)
gl_get_tex_image(GL_TEXTURE_2D, 0, GL_RGBA, GL_UNSIGNED_BYTE, buf)
let hdr = `P6\n{w} {h}\n255\n`
file_write(f, hdr, len(hdr))
let row = bytes(w * 3)
for y in 0 .. h {
for x in 0 .. w {
row[x * 3] = buf[(y * w + x) * 4]; row[x * 3 + 1] = buf[(y * w + x) * 4 + 1]; row[x * 3 + 2] = buf[(y * w + x) * 4 + 2]
if tex_dump_alpha { let a = buf[(y * w + x) * 4 + 3]; row[x * 3] = a; row[x * 3 + 1] = a; row[x * 3 + 2] = a }
}
file_write(f, row, w * 3)
}
file_close(f)
free(buf); free(row)
}
# A binary PPM (P6, what Gl.screenshot writes) as an RGB8 texture, box-filtered down by
# `shrink` (a photo thumbnail); 0 when the file is missing.
function tex_load_ppm(path: pointer, shrink: int) -> int {
let d = r3d_read_file(path)
if d == null { return 0 }
let size = tex_file_len
var i = 2
var w = 0; var h = 0; var mx = 0
var field = 0
while i < size and field < 3 {
while i < size and (d[i] == 32 or d[i] == 10 or d[i] == 13 or d[i] == 9) { i += 1 }
var v = 0
while i < size and d[i] >= '0' and d[i] <= '9' { v = v * 10 + (d[i] - 48); i += 1 }
if field == 0 { w = v } else if field == 1 { h = v } else { mx = v }
field += 1
}
i += 1
if w <= 0 or h <= 0 or i + w * h * 3 > size { free(d); return 0 }
var k = shrink
if k < 1 { k = 1 }
let ow = w / k; let oh = h / k
let px = bytes(ow * oh * 3)
for y in 0 .. oh {
for x in 0 .. ow {
var r = 0; var g = 0; var b = 0
for yy in 0 .. k { for xx in 0 .. k {
let o = i + ((y * k + yy) * w + x * k + xx) * 3
r += d[o]; g += d[o + 1]; b += d[o + 2]
} }
let n = k * k
let q = (y * ow + x) * 3
px[q] = r / n; px[q + 1] = g / n; px[q + 2] = b / n
}
}
free(d)
tex_w = ow; tex_h = oh; tex_channels = 3; tex_depth = 8
let id = tex_upload(px, true, false)
free(px)
return id
}

View file

@ -0,0 +1,125 @@
# ============================================================================
# water.ludic — still water for the valley floor: a level plane over a region,
# drawn after the opaque scene, showing only where the ground lies below it.
# Sky reflection with fresnel, sun glitter, scrolling ripple normals, a depth
# tinted body read from the scene depth, and soft shores.
# ============================================================================
var water_mesh: Mesh = null
var water_prog: int = 0
var water_level: int = 0
var water_cx: int = 0
var water_cz: int = 0
var water_ex: int = 0
var water_ez: int = 0
var water_on: bool = false
var water_refl: Target = null # the world mirrored in the surface, half resolution
var water_refl_div: int = 2 # R3D_REFLDIV overrides: 2 = half res, 4 = quarter
var water_saved: words = null # the real camera's matrices, restored after the pass
var water_dumped: bool = false
# Render the scene through a camera mirrored in the water plane into water_refl,
# clipping everything below the surface; terrain, scattered layers and sky.
function water_reflection_pass() -> void {
if water_refl == null {
if Os.has_env("R3D_REFLDIV") { water_refl_div = Text.to_int(Os.env("R3D_REFLDIV")) }
if water_refl != null { target_free(water_refl) }
water_refl = target_new(gl_w / water_refl_div, gl_h / water_refl_div, GL_RGBA16F, GL_RGBA, GL_HALF_FLOAT, true, GL_LINEAR)
water_saved = words(16 * 4 + 3)
}
# save the camera
m4_copy(water_saved, cam_view)
m4_copy(mem_off(water_saved, 64), cam_vp)
m4_copy(mem_off(water_saved, 128), cam_inv_vp)
let sx = cam_pos[0]; let sy = cam_pos[1]; let sz = cam_pos[2]
# the mirrored camera: view' = view * R, R reflecting y about the surface (y' = 2L - y).
# R has determinant -1, so the winding flips (front faces culled below) and the image
# lands exactly where the main camera's pixels expect the reflection.
let eye = v3_new(sx, f_sub(f_mul(F_TWO, water_level), sy), sz)
let refl = m4_new()
refl[5] = f_neg1()
refl[13] = f_mul(F_TWO, water_level)
let mv = words(16)
m4_mul(mv, water_saved, refl)
m4_copy(cam_view, mv)
free(mv); free(refl)
let fwd = words(3); let up = words(3); let at = words(3)
m4_mul(cam_vp, cam_proj, cam_view)
m4_inverse(cam_inv_vp, cam_vp)
v3_copy(cam_pos, eye)
r3d_clip_y = f_sub(water_level, fl(0.05))
target_bind(water_refl)
gl_enable(GL_DEPTH_TEST)
gl_depth_func(GL_LESS)
gl_depth_mask(1)
gl_enable(GL_CULL_FACE)
gl_cull_face(GL_FRONT) # the mirror flips the winding
gl_clear_color(0.0, 0.0, 0.0, 1.0)
gl_clear(GL_COLOR_BUFFER_BIT | GL_DEPTH_BUFFER_BIT)
sc_freeze = true
let sb = sc_skip_blade
sc_skip_blade = true # blades are invisible at this scale in a reflection
ter_reflect = true
terrain_draw()
scene_draw()
r3d_draw_sky()
ter_reflect = false
sc_skip_blade = sb
sc_freeze = false
gl_cull_face(GL_BACK)
if Os.has_env("R3D_DUMP_REFL") and not water_dumped { water_dumped = true; tex_dump(water_refl.color, gl_w / water_refl_div, gl_h / water_refl_div, "build/dbg_refl.ppm") }
# restore
r3d_clip_y = 0xCF000000
m4_copy(cam_view, water_saved)
m4_copy(cam_vp, mem_off(water_saved, 64))
m4_copy(cam_inv_vp, mem_off(water_saved, 128))
v3_set(cam_pos, sx, sy, sz)
free(eye); free(fwd); free(up); free(at)
gl_bind_framebuffer(GL_FRAMEBUFFER, 0)
}
function water_init(level: int, cx: int, cz: int, ex: int, ez: int) -> void {
water_mesh = mesh_grid(2, F_HALF)
water_prog = r3d_program("water.vert", "water.frag", "")
water_level = level; water_cx = cx; water_cz = cz; water_ex = ex; water_ez = ez
water_on = true
}
# call after the opaque pass, before the sky: blends over the resolved depth
function water_draw(depth_tex: int) -> void {
if not water_on { return }
let p = water_prog
gl_use_program(p)
u_mat4(gl_uniform(p, "u_view"), cam_view)
u_mat4(gl_uniform(p, "u_proj"), cam_proj)
u_mat4(gl_uniform(p, "u_inv_vp"), cam_inv_vp)
u_f(gl_uniform(p, "u_level"), water_level)
u_f2(gl_uniform(p, "u_center"), water_cx, water_cz)
u_f2(gl_uniform(p, "u_extent"), water_ex, water_ez)
u_f2(gl_uniform(p, "u_screen"), fi(gl_w), fi(gl_h))
var ron = F_ZERO
if water_refl != null {
# bind on its own unit first: generating the mip chain re-binds the texture on the active unit,
# and it must not displace the depth texture the shader reads for the shore
r3d_bind_2d(p, "u_refl", 1, water_refl.color); ron = F_ONE
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_MIN_FILTER, GL_LINEAR_MIPMAP_LINEAR)
gl_generate_mipmap(GL_TEXTURE_2D)
}
r3d_bind_2d(p, "u_depth", 0, depth_tex)
r3d_bind_2d(p, "u_scene", 2, post_scene.color)
u_f(gl_uniform(p, "u_refl_on"), ron)
sky_bind_lighting(p)
shadow_bind(p)
fog_bind(p)
# Opaque. The surface composites the refracted bed itself, so there is nothing for
# hardware blending to do — and an alpha was what left see-through gaps in the foam
# and a clear band at the shore wide enough to give the plane away.
gl_disable(GL_BLEND)
gl_blend_func(GL_SRC_ALPHA, GL_ONE_MINUS_SRC_ALPHA)
# the surface writes depth: the ambient-occlusion and temporal passes read the frame's depth,
# and the bed 9 m below the shore would otherwise darken a band along the water line
gl_depth_mask(1)
gl_disable(GL_CULL_FACE)
mesh_draw(water_mesh)
gl_disable(GL_BLEND)
}

View file

@ -81,6 +81,7 @@ declare i32 @CGWarpMouseCursorPosition(%NSPoint)
@.s_type = private unnamed_addr constant [5 x i8] c"type\00"
@.s_keycd = private unnamed_addr constant [8 x i8] c"keyCode\00"
@.s_chars = private unnamed_addr constant [28 x i8] c"charactersIgnoringModifiers\00"
@.s_modf = private unnamed_addr constant [14 x i8] c"modifierFlags\00"
@.s_length = private unnamed_addr constant [7 x i8] c"length\00"
@.s_charat = private unnamed_addr constant [18 x i8] c"characterAtIndex:\00"
@.s_locwin = private unnamed_addr constant [17 x i8] c"locationInWindow\00"
@ -267,9 +268,12 @@ entry:
%sel_alloc = call ptr @sel_registerName(ptr @.s_alloc)
%w0 = call ptr @objc_msgSend(ptr %wincls, ptr %sel_alloc)
%sel_initw = call ptr @sel_registerName(ptr @.s_initw)
; NSWindowStyleMaskTitled|Closable = 3, NSBackingStoreBuffered = 2
%win = call ptr (ptr, ptr, %CGRect, i64, i64, i8) @objc_msgSend(ptr %w0, ptr %sel_initw, %CGRect %rect, i64 3, i64 2, i8 0)
; NSWindowStyleMaskTitled|Closable|Resizable = 11, NSBackingStoreBuffered = 2
%win = call ptr (ptr, ptr, %CGRect, i64, i64, i8) @objc_msgSend(ptr %w0, ptr %sel_initw, %CGRect %rect, i64 11, i64 2, i8 0)
store ptr %win, ptr @W_win
; NSWindowCollectionBehaviorFullScreenPrimary = 1 << 7: the green button and toggleFullScreen: work
%sel_cb = call ptr @sel_registerName(ptr @.s_setcb)
%acb = call ptr (ptr, ptr, i64) @objc_msgSend(ptr %win, ptr %sel_cb, i64 128)
%strcls = call ptr @objc_getClass(ptr @.c_str)
%sel_utf8 = call ptr @sel_registerName(ptr @.s_utf8)
@ -343,10 +347,10 @@ entry:
i32 36, label %vret
i32 53, label %vesc
]
vw: ret i32 119
vs: ret i32 115
va: ret i32 97
vd: ret i32 100
vw: ret i32 128
vs: ret i32 129
va: ret i32 130
vd: ret i32 131
vspace: ret i32 32
vret: ret i32 10
vesc: ret i32 27
@ -361,7 +365,15 @@ chars:
take:
%c = call i16 (ptr, ptr, i64) @objc_msgSend(ptr %s, ptr %sel_cat, i64 0)
%c32 = zext i16 %c to i32
ret i32 %c32
; charactersIgnoringModifiers still honours Shift, so Shift+W arrives as 'W' and a game
; holding 'w' to walk stopped dead the moment the player held Shift to run. Fold A-Z to
; a-z: the held set is keyed by the key, Shift itself is reported separately (code 16).
%isupA = icmp sge i32 %c32, 65
%isupB = icmp sle i32 %c32, 90
%isupper = and i1 %isupA, %isupB
%lower = add i32 %c32, 32
%folded = select i1 %isupper, i32 %lower, i32 %c32
ret i32 %folded
none:
ret i32 0
}
@ -401,11 +413,30 @@ handle:
br i1 %iskey, label %key, label %notkey
notkey:
%isup = icmp eq i64 %ty, 11 ; NSEventTypeKeyUp
br i1 %isup, label %keyup, label %mouse
br i1 %isup, label %keyup, label %notup
keyup: ; #50 — release the held key
%uv = call i32 @ev_keyval(ptr %ev)
call void @win_held_bit(i32 %uv, i32 0)
br label %forward
notup:
%isflags = icmp eq i64 %ty, 12 ; NSEventTypeFlagsChanged: modifier keys
br i1 %isflags, label %flags, label %mouse
flags: ; Shift is held-key 16 (ctrl 17, alt 18): a run modifier
%sel_mf = call ptr @sel_registerName(ptr @.s_modf)
%mf = call i64 (ptr, ptr) @objc_msgSend(ptr %ev, ptr %sel_mf)
%mf_sh = and i64 %mf, 131072 ; NSEventModifierFlagShift = 1 << 17
%sh_on = icmp ne i64 %mf_sh, 0
%sh_i = zext i1 %sh_on to i32
call void @win_held_bit(i32 16, i32 %sh_i)
%mf_ct = and i64 %mf, 262144 ; NSEventModifierFlagControl = 1 << 18
%ct_on = icmp ne i64 %mf_ct, 0
%ct_i = zext i1 %ct_on to i32
call void @win_held_bit(i32 17, i32 %ct_i)
%mf_al = and i64 %mf, 524288 ; NSEventModifierFlagOption = 1 << 19
%al_on = icmp ne i64 %mf_al, 0
%al_i = zext i1 %al_on to i32
call void @win_held_bit(i32 18, i32 %al_i)
br label %forward
mouse: ; #50 — mouse buttons + wheel
%ml_d = icmp eq i64 %ty, 1 ; NSEventTypeLeftMouseDown
br i1 %ml_d, label %lset, label %ml_u
@ -610,8 +641,17 @@ reldelta:
%nmy2 = select i1 %yhi, i32 %fbhm, i32 %nmy1
store i32 %nmx2, ptr @W_mx
store i32 %nmy2, ptr @W_my
; the raw motion too: a clamped cursor stops turning at the edge, the delta must not
%p4r = getelementptr i32, ptr %out, i32 4
store i32 %ddx, ptr %p4r
%p5r = getelementptr i32, ptr %out, i32 5
store i32 %ddy, ptr %p5r
br label %emit
abspos:
%p4a = getelementptr i32, ptr %out, i32 4
store i32 0, ptr %p4a
%p5a = getelementptr i32, ptr %out, i32 5
store i32 0, ptr %p5a
%win = load ptr, ptr @W_win
%nowin = icmp eq ptr %win, null
br i1 %nowin, label %emit, label %qpos
@ -1107,3 +1147,264 @@ body:
ret:
ret void
}
; ============================================================================
; OpenGL on the window (Gl.* — runtime/native/gl.ludic). An NSOpenGLContext is
; attached to the existing LudicView, at the display's backing resolution, with
; a 4.1 core profile. The CPU framebuffer path above is untouched: a Gl program
; simply never calls Screen.show, so -drawRect: finds @W_fb null and paints
; nothing. The window-independent GL pieces (offscreen CGL contexts, ABI thunks)
; live in gl.ll so a headless build never references the window.
;
; win_gl_attach() -> ok create the context on @W_view (0 = no window yet)
; win_gl_resize(w, h) content size in points, scale 1 (mouse mapping follows)
; win_gl_swap() flushBuffer (vsync'd)
; win_gl_scale() -> int backing pixels per point (2 on Retina)
; ============================================================================
@.c_pixfmt = private unnamed_addr constant [20 x i8] c"NSOpenGLPixelFormat\00"
@.c_glctx = private unnamed_addr constant [16 x i8] c"NSOpenGLContext\00"
@.s_initattr = private unnamed_addr constant [20 x i8] c"initWithAttributes:\00"
@.s_initfmt = private unnamed_addr constant [29 x i8] c"initWithFormat:shareContext:\00"
@.s_setview = private unnamed_addr constant [9 x i8] c"setView:\00"
@.s_makecur = private unnamed_addr constant [19 x i8] c"makeCurrentContext\00"
@.s_flushbuf = private unnamed_addr constant [12 x i8] c"flushBuffer\00"
@.s_ctxupd = private unnamed_addr constant [7 x i8] c"update\00"
@.s_setvals = private unnamed_addr constant [24 x i8] c"setValues:forParameter:\00"
@.s_bestres = private unnamed_addr constant [37 x i8] c"setWantsBestResolutionOpenGLSurface:\00"
@.s_bscale = private unnamed_addr constant [19 x i8] c"backingScaleFactor\00"
@.s_setcsize = private unnamed_addr constant [16 x i8] c"setContentSize:\00"
@.s_togfs = private unnamed_addr constant [18 x i8] c"toggleFullScreen:\00"
@.s_bounds = private unnamed_addr constant [7 x i8] c"bounds\00"
@.s_setcb = private unnamed_addr constant [23 x i8] c"setCollectionBehavior:\00"
@W_glctx = internal global ptr null
@W_glscale = internal global i32 1
define i32 @win_gl_attach() {
entry:
%view = load ptr, ptr @W_view
%noview = icmp eq ptr %view, null
br i1 %noview, label %fail, label %go
go:
%sel_alloc = call ptr @sel_registerName(ptr @.s_alloc)
; Retina: ask for a backing-resolution surface before the context is attached
%sel_br = call ptr @sel_registerName(ptr @.s_bestres)
%r0 = call ptr (ptr, ptr, i8) @objc_msgSend(ptr %view, ptr %sel_br, i8 1)
; NSOpenGLPixelFormatAttribute[]: OpenGLProfile=99 -> 4.1 core (0x4100),
; ColorSize=8 -> 24, AlphaSize=11 -> 8, DepthSize=12 -> 24, StencilSize=13 -> 8,
; DoubleBuffer=5, Accelerated=73, 0
%attrs = alloca [14 x i32], align 4
%a0 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 0
store i32 99, ptr %a0
%a1 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 1
store i32 16640, ptr %a1
%a2 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 2
store i32 8, ptr %a2
%a3 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 3
store i32 24, ptr %a3
%a4 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 4
store i32 11, ptr %a4
%a5 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 5
store i32 8, ptr %a5
%a6 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 6
store i32 12, ptr %a6
%a7 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 7
store i32 24, ptr %a7
%a8 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 8
store i32 13, ptr %a8
%a9 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 9
store i32 8, ptr %a9
%a10 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 10
store i32 5, ptr %a10
%a11 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 11
store i32 73, ptr %a11
%a12 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 12
store i32 0, ptr %a12
%a13 = getelementptr [14 x i32], ptr %attrs, i32 0, i32 13
store i32 0, ptr %a13
%pfcls = call ptr @objc_getClass(ptr @.c_pixfmt)
%pf0 = call ptr @objc_msgSend(ptr %pfcls, ptr %sel_alloc)
%sel_ia = call ptr @sel_registerName(ptr @.s_initattr)
%pf = call ptr (ptr, ptr, ptr) @objc_msgSend(ptr %pf0, ptr %sel_ia, ptr %attrs)
%nopf = icmp eq ptr %pf, null
br i1 %nopf, label %fail, label %ctx
ctx:
%ccls = call ptr @objc_getClass(ptr @.c_glctx)
%c0 = call ptr @objc_msgSend(ptr %ccls, ptr %sel_alloc)
%sel_if = call ptr @sel_registerName(ptr @.s_initfmt)
%c = call ptr (ptr, ptr, ptr, ptr) @objc_msgSend(ptr %c0, ptr %sel_if, ptr %pf, ptr null)
%noc = icmp eq ptr %c, null
br i1 %noc, label %fail, label %attach
attach:
%sel_sv = call ptr @sel_registerName(ptr @.s_setview)
%r1 = call ptr (ptr, ptr, ptr) @objc_msgSend(ptr %c, ptr %sel_sv, ptr %view)
%sel_mc = call ptr @sel_registerName(ptr @.s_makecur)
%r2 = call ptr @objc_msgSend(ptr %c, ptr %sel_mc)
; swap interval 1 (NSOpenGLCPSwapInterval = 222)
%one = alloca i32, align 4
store i32 1, ptr %one
%sel_sp = call ptr @sel_registerName(ptr @.s_setvals)
%r3 = call ptr (ptr, ptr, ptr, i64) @objc_msgSend(ptr %c, ptr %sel_sp, ptr %one, i64 222)
store ptr %c, ptr @W_glctx
; backing scale factor of the window (2.0 on Retina)
%win = load ptr, ptr @W_win
%sel_bs = call ptr @sel_registerName(ptr @.s_bscale)
%bs = call double (ptr, ptr) @objc_msgSend(ptr %win, ptr %sel_bs)
%bsi = fptosi double %bs to i32
%bsok = icmp sgt i32 %bsi, 0
%bsv = select i1 %bsok, i32 %bsi, i32 1
store i32 %bsv, ptr @W_glscale
ret i32 1
fail:
ret i32 0
}
define void @win_gl_resize(i32 %w, i32 %h) {
entry:
%win = load ptr, ptr @W_win
%nowin = icmp eq ptr %win, null
br i1 %nowin, label %out, label %go
go:
store i32 %w, ptr @W_fbw
store i32 %h, ptr @W_fbh
store i32 1, ptr @W_scale
%wd = sitofp i32 %w to double
%hd = sitofp i32 %h to double
%s0 = insertvalue %NSPoint undef, double %wd, 0
%sz = insertvalue %NSPoint %s0, double %hd, 1
%sel_cs = call ptr @sel_registerName(ptr @.s_setcsize)
%r0 = call ptr (ptr, ptr, %NSPoint) @objc_msgSend(ptr %win, ptr %sel_cs, %NSPoint %sz)
%sel_center = call ptr @sel_registerName(ptr @.s_center)
%r1 = call ptr @objc_msgSend(ptr %win, ptr %sel_center)
%c = load ptr, ptr @W_glctx
%noc = icmp eq ptr %c, null
br i1 %noc, label %out, label %upd
upd:
%sel_u = call ptr @sel_registerName(ptr @.s_ctxupd)
%r2 = call ptr @objc_msgSend(ptr %c, ptr %sel_u)
br label %out
out:
ret void
}
define void @win_gl_swap() {
entry:
%c = load ptr, ptr @W_glctx
%noc = icmp eq ptr %c, null
br i1 %noc, label %out, label %go
go:
%sel_fb = call ptr @sel_registerName(ptr @.s_flushbuf)
%r = call ptr @objc_msgSend(ptr %c, ptr %sel_fb)
br label %out
out:
ret void
}
define i32 @win_gl_scale() {
entry:
%s = load i32, ptr @W_glscale
ret i32 %s
}
; swap interval: 1 = vsync (the default set at attach), 0 = free-running
define void @win_gl_swap_interval(i32 %n) {
entry:
%c = load ptr, ptr @W_glctx
%noc = icmp eq ptr %c, null
br i1 %noc, label %out, label %go
go:
%v = alloca i32, align 4
store i32 %n, ptr %v
%sel_sp = call ptr @sel_registerName(ptr @.s_setvals)
%r = call ptr (ptr, ptr, ptr, i64) @objc_msgSend(ptr %c, ptr %sel_sp, ptr %v, i64 222)
br label %out
out:
ret void
}
; ============================================================================
; window size, full screen and the drawable (the game's video settings)
; ============================================================================
; the view's drawable in pixels -> out[0], out[1]; also keeps the points size the
; mouse conversion uses in step with a window the user resized or made full screen
define void @win_gl_drawable(ptr %out) {
entry:
%view = load ptr, ptr @W_view
%noview = icmp eq ptr %view, null
br i1 %noview, label %none, label %go
go:
%sel_b = call ptr @sel_registerName(ptr @.s_bounds)
%r = call %CGRect (ptr, ptr) @objc_msgSend(ptr %view, ptr %sel_b)
%w = extractvalue %CGRect %r, 2
%h = extractvalue %CGRect %r, 3
%sc = load i32, ptr @W_glscale
%scd = sitofp i32 %sc to double
%pw = fmul double %w, %scd
%ph = fmul double %h, %scd
%iw = fptosi double %pw to i32
%ih = fptosi double %ph to i32
store i32 %iw, ptr %out
%p1 = getelementptr i32, ptr %out, i32 1
store i32 %ih, ptr %p1
%ipw = fptosi double %w to i32
%iph = fptosi double %h to i32
store i32 %ipw, ptr @W_fbw
store i32 %iph, ptr @W_fbh
ret void
none:
store i32 0, ptr %out
%p1n = getelementptr i32, ptr %out, i32 1
store i32 0, ptr %p1n
ret void
}
; tell the context its view changed size
define void @win_gl_update() {
entry:
%c = load ptr, ptr @W_glctx
%noc = icmp eq ptr %c, null
br i1 %noc, label %out, label %upd
upd:
%sel_u = call ptr @sel_registerName(ptr @.s_ctxupd)
%r = call ptr @objc_msgSend(ptr %c, ptr %sel_u)
br label %out
out:
ret void
}
define void @win_toggle_fullscreen() {
entry:
%win = load ptr, ptr @W_win
%nowin = icmp eq ptr %win, null
br i1 %nowin, label %out, label %go
go:
%sel_t = call ptr @sel_registerName(ptr @.s_togfs)
%r = call ptr (ptr, ptr, ptr) @objc_msgSend(ptr %win, ptr %sel_t, ptr null)
br label %out
out:
ret void
}
; Retina (backing-resolution) drawable on or off: the pixel scale follows
define void @win_gl_retina(i32 %on) {
entry:
%view = load ptr, ptr @W_view
%noview = icmp eq ptr %view, null
br i1 %noview, label %out, label %go
go:
%sel_br = call ptr @sel_registerName(ptr @.s_bestres)
%flag = trunc i32 %on to i8
%r0 = call ptr (ptr, ptr, i8) @objc_msgSend(ptr %view, ptr %sel_br, i8 %flag)
%win = load ptr, ptr @W_win
%sel_bs = call ptr @sel_registerName(ptr @.s_bscale)
%bs = call double (ptr, ptr) @objc_msgSend(ptr %win, ptr %sel_bs)
%bsi = fptosi double %bs to i32
%bsok = icmp sgt i32 %bsi, 0
%bsv = select i1 %bsok, i32 %bsi, i32 1
%ison = icmp ne i32 %on, 0
%scale = select i1 %ison, i32 %bsv, i32 1
store i32 %scale, ptr @W_glscale
call void @win_gl_update()
br label %out
out:
ret void
}

368
runtime/native/gl.ll Normal file
View file

@ -0,0 +1,368 @@
; ============================================================================
; gl.ll — the window-independent half of the OpenGL backend, in LLVM IR.
;
; Linked into any program that uses Gl.* (windowed or headless), together with
; gl_thunks.ll (the generated per-entry-point ABI thunks) and -framework OpenGL.
; Nothing here touches the window: the NSOpenGLContext lives in cocoa.ll.
;
; cgl_offscreen() -> ok a headless 4.1 core context (render into FBOs)
; fx_to_f32(fx) -> bits Q16.16 -> IEEE float bits (an int)
; f32_to_fx(bits) -> fx IEEE float bits -> Q16.16
; mem_* raw little-endian reads/writes on a bytes buffer
; f_* IEEE-754 float arithmetic on float bits
; ============================================================================
declare i32 @CGLChoosePixelFormat(ptr, ptr, ptr)
declare i32 @CGLCreateContext(ptr, ptr, ptr)
declare i32 @CGLSetCurrentContext(ptr)
declare i32 @CGLDestroyPixelFormat(ptr)
declare void @llvm.memcpy.p0.p0.i64(ptr, ptr, i64, i1)
declare void @llvm.memset.p0.i64(ptr, i8, i64, i1)
declare float @sinf(float)
declare float @cosf(float)
declare float @tanf(float)
declare float @atan2f(float, float)
declare float @powf(float, float)
declare float @expf(float)
declare float @logf(float)
declare float @floorf(float)
declare float @fmodf(float, float)
declare float @ldexpf(float, i32)
declare float @llvm.sqrt.f32(float)
declare float @llvm.fabs.f32(float)
define i32 @cgl_offscreen() {
entry:
; kCGLPFAAccelerated=73, kCGLPFAOpenGLProfile=99 -> kCGLOGLPVersion_GL4_Core (0x4100),
; kCGLPFAColorSize=8 -> 24, kCGLPFADepthSize=12 -> 24, 0
%attrs = alloca [8 x i32], align 4
%a0 = getelementptr [8 x i32], ptr %attrs, i32 0, i32 0
store i32 73, ptr %a0
%a1 = getelementptr [8 x i32], ptr %attrs, i32 0, i32 1
store i32 99, ptr %a1
%a2 = getelementptr [8 x i32], ptr %attrs, i32 0, i32 2
store i32 16640, ptr %a2
%a3 = getelementptr [8 x i32], ptr %attrs, i32 0, i32 3
store i32 8, ptr %a3
%a4 = getelementptr [8 x i32], ptr %attrs, i32 0, i32 4
store i32 24, ptr %a4
%a5 = getelementptr [8 x i32], ptr %attrs, i32 0, i32 5
store i32 12, ptr %a5
%a6 = getelementptr [8 x i32], ptr %attrs, i32 0, i32 6
store i32 24, ptr %a6
%a7 = getelementptr [8 x i32], ptr %attrs, i32 0, i32 7
store i32 0, ptr %a7
%pix = alloca ptr, align 8
store ptr null, ptr %pix
%npix = alloca i32, align 4
%e1 = call i32 @CGLChoosePixelFormat(ptr %attrs, ptr %pix, ptr %npix)
%p = load ptr, ptr %pix
%nop = icmp eq ptr %p, null
br i1 %nop, label %fail, label %mk
mk:
%ctx = alloca ptr, align 8
store ptr null, ptr %ctx
%e2 = call i32 @CGLCreateContext(ptr %p, ptr null, ptr %ctx)
%c = load ptr, ptr %ctx
%e3 = call i32 @CGLDestroyPixelFormat(ptr %p)
%noc = icmp eq ptr %c, null
br i1 %noc, label %fail, label %cur
cur:
%e4 = call i32 @CGLSetCurrentContext(ptr %c)
ret i32 1
fail:
ret i32 0
}
; ---- Q16.16 <-> IEEE float ---------------------------------------------------
define i32 @fx_to_f32(i32 %fx) {
entry:
%f = sitofp i32 %fx to float
%s = fmul float %f, 0x3EF0000000000000
%b = bitcast float %s to i32
ret i32 %b
}
define i32 @f32_to_fx(i32 %bits) {
entry:
%f = bitcast i32 %bits to float
%s = fmul float %f, 65536.0
%r = fptosi float %s to i32
ret i32 %r
}
; ---- raw memory ---------------------------------------------------------------
define ptr @mem_off(ptr %p, i32 %off) {
entry:
%q = getelementptr inbounds i8, ptr %p, i32 %off
ret ptr %q
}
define i32 @mem_get_i32(ptr %p, i32 %off) {
entry:
%q = getelementptr inbounds i8, ptr %p, i32 %off
%v = load i32, ptr %q, align 1
ret i32 %v
}
define void @mem_put_i32(ptr %p, i32 %off, i32 %v) {
entry:
%q = getelementptr inbounds i8, ptr %p, i32 %off
store i32 %v, ptr %q, align 1
ret void
}
define i32 @mem_get_u16(ptr %p, i32 %off) {
entry:
%q = getelementptr inbounds i8, ptr %p, i32 %off
%v = load i16, ptr %q, align 1
%z = zext i16 %v to i32
ret i32 %z
}
define void @mem_put_u16(ptr %p, i32 %off, i32 %v) {
entry:
%q = getelementptr inbounds i8, ptr %p, i32 %off
%t = trunc i32 %v to i16
store i16 %t, ptr %q, align 1
ret void
}
define i32 @mem_get_u8(ptr %p, i32 %off) {
entry:
%q = getelementptr inbounds i8, ptr %p, i32 %off
%v = load i8, ptr %q
%z = zext i8 %v to i32
ret i32 %z
}
define void @mem_put_u8(ptr %p, i32 %off, i32 %v) {
entry:
%q = getelementptr inbounds i8, ptr %p, i32 %off
%t = trunc i32 %v to i8
store i8 %t, ptr %q
ret void
}
; float element i of a float buffer, as Q16.16 / as bits
define i32 @mem_get_f32(ptr %p, i32 %i) {
entry:
%q = getelementptr inbounds float, ptr %p, i32 %i
%f = load float, ptr %q, align 1
%s = fmul float %f, 65536.0
%r = fptosi float %s to i32
ret i32 %r
}
define void @mem_put_f32(ptr %p, i32 %i, i32 %fx) {
entry:
%q = getelementptr inbounds float, ptr %p, i32 %i
%f = sitofp i32 %fx to float
%s = fmul float %f, 0x3EF0000000000000
store float %s, ptr %q, align 1
ret void
}
define i32 @mem_get_f32_bits(ptr %p, i32 %i) {
entry:
%q = getelementptr inbounds i32, ptr %p, i32 %i
%v = load i32, ptr %q, align 1
ret i32 %v
}
define void @mem_put_f32_bits(ptr %p, i32 %i, i32 %bits) {
entry:
%q = getelementptr inbounds i32, ptr %p, i32 %i
store i32 %bits, ptr %q, align 1
ret void
}
define void @mem_copy(ptr %dst, ptr %src, i32 %n) {
entry:
%n64 = sext i32 %n to i64
call void @llvm.memcpy.p0.p0.i64(ptr %dst, ptr %src, i64 %n64, i1 false)
ret void
}
define void @mem_set(ptr %dst, i32 %v, i32 %n) {
entry:
%n64 = sext i32 %n to i64
%b = trunc i32 %v to i8
call void @llvm.memset.p0.i64(ptr %dst, i8 %b, i64 %n64, i1 false)
ret void
}
; ---- IEEE float arithmetic on bit patterns ------------------------------------
define i32 @f_add(i32 %a, i32 %b) {
entry:
%x = bitcast i32 %a to float
%y = bitcast i32 %b to float
%r = fadd float %x, %y
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_sub(i32 %a, i32 %b) {
entry:
%x = bitcast i32 %a to float
%y = bitcast i32 %b to float
%r = fsub float %x, %y
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_mul(i32 %a, i32 %b) {
entry:
%x = bitcast i32 %a to float
%y = bitcast i32 %b to float
%r = fmul float %x, %y
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_div(i32 %a, i32 %b) {
entry:
%x = bitcast i32 %a to float
%y = bitcast i32 %b to float
%r = fdiv float %x, %y
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_neg(i32 %a) {
entry:
%x = bitcast i32 %a to float
%r = fneg float %x
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_sqrt(i32 %a) {
entry:
%x = bitcast i32 %a to float
%r = call float @llvm.sqrt.f32(float %x)
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_abs(i32 %a) {
entry:
%x = bitcast i32 %a to float
%r = call float @llvm.fabs.f32(float %x)
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_sin(i32 %a) {
entry:
%x = bitcast i32 %a to float
%r = call float @sinf(float %x)
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_cos(i32 %a) {
entry:
%x = bitcast i32 %a to float
%r = call float @cosf(float %x)
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_tan(i32 %a) {
entry:
%x = bitcast i32 %a to float
%r = call float @tanf(float %x)
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_atan2(i32 %a, i32 %b) {
entry:
%x = bitcast i32 %a to float
%y = bitcast i32 %b to float
%r = call float @atan2f(float %x, float %y)
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_pow(i32 %a, i32 %b) {
entry:
%x = bitcast i32 %a to float
%y = bitcast i32 %b to float
%r = call float @powf(float %x, float %y)
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_exp(i32 %a) {
entry:
%x = bitcast i32 %a to float
%r = call float @expf(float %x)
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_log(i32 %a) {
entry:
%x = bitcast i32 %a to float
%r = call float @logf(float %x)
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_floor(i32 %a) {
entry:
%x = bitcast i32 %a to float
%r = call float @floorf(float %x)
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_mod(i32 %a, i32 %b) {
entry:
%x = bitcast i32 %a to float
%y = bitcast i32 %b to float
%r = call float @fmodf(float %x, float %y)
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_ldexp(i32 %a, i32 %e) {
entry:
%x = bitcast i32 %a to float
%r = call float @ldexpf(float %x, i32 %e)
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_min(i32 %a, i32 %b) {
entry:
%x = bitcast i32 %a to float
%y = bitcast i32 %b to float
%c = fcmp olt float %x, %y
%r = select i1 %c, float %x, float %y
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_max(i32 %a, i32 %b) {
entry:
%x = bitcast i32 %a to float
%y = bitcast i32 %b to float
%c = fcmp ogt float %x, %y
%r = select i1 %c, float %x, float %y
%o = bitcast float %r to i32
ret i32 %o
}
define i32 @f_lt(i32 %a, i32 %b) {
entry:
%x = bitcast i32 %a to float
%y = bitcast i32 %b to float
%c = fcmp olt float %x, %y
%z = zext i1 %c to i32
ret i32 %z
}
define i32 @f_from_int(i32 %a) {
entry:
%f = sitofp i32 %a to float
%o = bitcast float %f to i32
ret i32 %o
}
define i32 @f_to_int(i32 %a) {
entry:
%x = bitcast i32 %a to float
%r = fptosi float %x to i32
ret i32 %r
}
; ---- a real monotonic microsecond clock -------------------------------------
; Time.now() is whole seconds and Time.delta() is a fixed 60 Hz timestep, so
; neither can measure a frame. gettimeofday gives the wall clock in microseconds,
; which is what frame pacing and hitch measurement actually need.
; macOS arm64: struct timeval is { time_t tv_sec (i64), suseconds_t tv_usec (i32) }
%struct.timeval64 = type { i64, i32 }
declare i32 @gettimeofday(ptr, ptr)
define i64 @gl_now_us() {
entry:
%tv = alloca %struct.timeval64, align 8
%r = call i32 @gettimeofday(ptr %tv, ptr null)
%secp = getelementptr inbounds %struct.timeval64, ptr %tv, i32 0, i32 0
%usp = getelementptr inbounds %struct.timeval64, ptr %tv, i32 0, i32 1
%sec = load i64, ptr %secp, align 8
%us32 = load i32, ptr %usp, align 8
%us = sext i32 %us32 to i64
%m = mul i64 %sec, 1000000
%t = add i64 %m, %us
ret i64 %t
}

282
runtime/native/gl.ludic Normal file
View file

@ -0,0 +1,282 @@
# ============================================================================
# gl.ludic — Gl.*: OpenGL for Ludic.
#
# The whole OpenGL 4.1 core API is available as Gl.<snake_name>(...) — every
# entry point of the platform gl3.h is bound in gl_api.ludic (generated), with
# every GL_* constant. float/double parameters take `fixed`; buffers are the raw
# `bytes`/`words` pointers Ludic already has, and pixel/vertex data is uploaded
# from them as-is. This file adds the small amount of glue the API needs to be
# usable from a game: a context on the window (or an offscreen one headless),
# the swap, a screenshot, shader/program helpers, and IEEE float helpers so a
# program can fill a vertex buffer with real floats from its Q16.16 math.
#
# Windowed: the NSOpenGLContext is attached to the existing LudicView (cocoa.ll)
# at the display's backing resolution. Headless: a CGL context with no drawable
# (gl.ll) and a framebuffer object that stands in for the screen, so the same
# program renders and screenshots byte-identically under the test harness.
# ============================================================================
import "gl_api.ludic"
# ---- native glue (cocoa.ll / gl.ll) -------------------------------------------
extern function win_gl_attach() -> int = "win_gl_attach"
extern function win_gl_resize(w: int, h: int) = "win_gl_resize"
extern function win_gl_swap() = "win_gl_swap"
extern function win_gl_scale() -> int = "win_gl_scale"
extern function win_gl_swap_interval(n: int) = "win_gl_swap_interval"
extern function win_gl_drawable(out: pointer) = "win_gl_drawable"
extern function win_gl_update() = "win_gl_update"
extern function win_gl_retina(on: int) = "win_gl_retina"
extern function win_toggle_fullscreen() = "win_toggle_fullscreen"
extern function cgl_offscreen() -> int = "cgl_offscreen"
# wall clock in microseconds — the only sub-second clock available to a Ludic program
extern function gl_now_us() -> long = "gl_now_us"
extern function fx_to_f32(fx: fixed) -> int = "fx_to_f32"
extern function f32_to_fx(bits: int) -> fixed = "f32_to_fx"
extern function mem_off(p: pointer, off: int) -> pointer = "mem_off"
extern function mem_get_i32(p: pointer, off: int) -> int = "mem_get_i32"
extern function mem_put_i32(p: pointer, off: int, v: int) = "mem_put_i32"
extern function mem_get_u16(p: pointer, off: int) -> int = "mem_get_u16"
extern function mem_put_u16(p: pointer, off: int, v: int) = "mem_put_u16"
extern function mem_get_u8(p: pointer, off: int) -> int = "mem_get_u8"
extern function mem_put_u8(p: pointer, off: int, v: int) = "mem_put_u8"
extern function mem_get_f32(p: pointer, i: int) -> fixed = "mem_get_f32"
extern function mem_put_f32(p: pointer, i: int, v: fixed) = "mem_put_f32"
extern function mem_get_f32_bits(p: pointer, i: int) -> int = "mem_get_f32_bits"
extern function mem_put_f32_bits(p: pointer, i: int, bits: int) = "mem_put_f32_bits"
extern function mem_copy(dst: pointer, src: pointer, n: int) = "mem_copy"
extern function mem_set(dst: pointer, v: int, n: int) = "mem_set"
# IEEE-754 single precision, carried as its bit pattern in an int
extern function f_add(a: int, b: int) -> int = "f_add"
extern function f_sub(a: int, b: int) -> int = "f_sub"
extern function f_mul(a: int, b: int) -> int = "f_mul"
extern function f_div(a: int, b: int) -> int = "f_div"
extern function f_neg(a: int) -> int = "f_neg"
extern function f_sqrt(a: int) -> int = "f_sqrt"
extern function f_abs(a: int) -> int = "f_abs"
extern function f_sin(a: int) -> int = "f_sin"
extern function f_cos(a: int) -> int = "f_cos"
extern function f_tan(a: int) -> int = "f_tan"
extern function f_atan2(a: int, b: int) -> int = "f_atan2"
extern function f_pow(a: int, b: int) -> int = "f_pow"
extern function f_exp(a: int) -> int = "f_exp"
extern function f_log(a: int) -> int = "f_log"
extern function f_floor(a: int) -> int = "f_floor"
extern function f_mod(a: int, b: int) -> int = "f_mod"
extern function f_ldexp(a: int, e: int) -> int = "f_ldexp"
extern function f_min(a: int, b: int) -> int = "f_min"
extern function f_max(a: int, b: int) -> int = "f_max"
extern function f_lt(a: int, b: int) -> int = "f_lt"
extern function f_from_int(a: int) -> int = "f_from_int"
extern function f_to_int(a: int) -> int = "f_to_int"
# ---- state --------------------------------------------------------------------
var gl_is_open: bool = false
var gl_w: int = 0 # drawable width, in pixels
var gl_h: int = 0
var gl_scale: int = 1 # backing pixels per window point
var gl_screen: int = 0 # the framebuffer that is "the screen" (an FBO headless)
var gl_ids: words = null # one-word scratch for glGen*/glGet*
function gl_scratch() -> words {
if gl_ids == null { gl_ids = words(4) }
return gl_ids
}
# Open a GL 4.1 core context on a w x h (points) window titled `title`; headless,
# an offscreen context with a w x h framebuffer standing in for the screen.
function gl_open(width: int, height: int, title: pointer) -> bool {
if gl_is_open { return true }
if is_windowed() {
if win_gl_attach() == 0 {
win_open(width, height, 1, title) # a plain program: no window yet
if win_gl_attach() == 0 { return false }
}
win_gl_resize(width, height)
gl_scale = win_gl_scale()
gl_w = width * gl_scale
gl_h = height * gl_scale
gl_screen = 0
} else {
if cgl_offscreen() == 0 { return false }
gl_scale = 1
gl_w = width
gl_h = height
gl_screen = gl_make_screen_fbo(width, height)
}
gl_bind_framebuffer(GL_FRAMEBUFFER, gl_screen)
gl_viewport(0, 0, gl_w, gl_h)
gl_is_open = true
return true
}
function gl_make_screen_fbo(w: int, h: int) -> int {
let ids = gl_scratch()
gl_gen_framebuffers(1, ids)
let fbo = ids[0]
gl_bind_framebuffer(GL_FRAMEBUFFER, fbo)
gl_gen_textures(1, ids)
let tex = ids[0]
gl_bind_texture(GL_TEXTURE_2D, tex)
gl_tex_image2d(GL_TEXTURE_2D, 0, GL_RGBA8, w, h, 0, GL_RGBA, GL_UNSIGNED_BYTE, null)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_MIN_FILTER, GL_LINEAR)
gl_tex_parameteri(GL_TEXTURE_2D, GL_TEXTURE_MAG_FILTER, GL_LINEAR)
gl_framebuffer_texture2d(GL_FRAMEBUFFER, GL_COLOR_ATTACHMENT0, GL_TEXTURE_2D, tex, 0)
gl_gen_renderbuffers(1, ids)
let rb = ids[0]
gl_bind_renderbuffer(GL_RENDERBUFFER, rb)
gl_renderbuffer_storage(GL_RENDERBUFFER, GL_DEPTH24_STENCIL8, w, h)
gl_framebuffer_renderbuffer(GL_FRAMEBUFFER, GL_DEPTH_STENCIL_ATTACHMENT, GL_RENDERBUFFER, rb)
return fbo
}
# Did the drawable change size (a window drag, full screen, a Retina switch)? Then
# gl_w / gl_h follow it and the caller rebuilds its screen-sized targets.
var gl_size_buf: words = null
function gl_resize_check() -> bool {
if not is_windowed() or not gl_is_open { return false }
if gl_size_buf == null { gl_size_buf = words(4) }
win_gl_drawable(gl_size_buf)
let w = gl_size_buf[0]; let h = gl_size_buf[1]
if w <= 0 or h <= 0 { return false }
if w == gl_w and h == gl_h { return false }
win_gl_update()
gl_w = w; gl_h = h
gl_scale = win_gl_scale()
gl_viewport(0, 0, gl_w, gl_h)
return true
}
function gl_set_window(w: int, h: int) -> void { if is_windowed() { win_gl_resize(w, h) } }
function gl_toggle_fullscreen() -> void { if is_windowed() { win_toggle_fullscreen() } }
function gl_set_retina(on: bool) -> void { if is_windowed() { var v = 0; if on { v = 1 }; win_gl_retina(v) } }
# vsync on (1, the default) or off (0); headless has nothing to sync to
function gl_vsync(n: int) -> void { if is_windowed() { win_gl_swap_interval(n) } }
function gl_width() -> int { return gl_w }
function gl_height() -> int { return gl_h }
function gl_screen_fbo() -> int { return gl_screen }
function gl_pixel_scale() -> int { return gl_scale }
# Present the frame (vsync'd flushBuffer); headless, just finish the GPU work.
function gl_swap() -> void {
if is_windowed() { win_gl_swap() }
else { gl_finish() }
}
# Write what is on the screen framebuffer to a binary PPM (call before Gl.swap).
function gl_screenshot(path: pointer) -> bool {
let w = gl_w
let h = gl_h
let f = file_open(path, "wb")
if f == null { return false }
let buf = bytes(w * h * 3)
gl_bind_framebuffer(GL_READ_FRAMEBUFFER, gl_screen)
gl_pixel_storei(GL_PACK_ALIGNMENT, 1)
gl_read_pixels(0, 0, w, h, GL_RGB, GL_UNSIGNED_BYTE, buf)
let hdr = `P6\n{w} {h}\n255\n`
file_write(f, hdr, len(hdr))
var y = h - 1
while y >= 0 {
file_write(f, mem_off(buf, y * w * 3), w * 3)
y -= 1
}
file_close(f)
free(buf)
return true
}
# Print any pending GL error under a tag; returns the error code (0 = none).
function gl_check(tag: pointer) -> int {
let e = gl_get_error()
if e != 0 { print(`gl error {e} at {tag}`) }
return e
}
# ---- shaders --------------------------------------------------------------------
var gl_log_buf: string = null
# Compile one shader stage from source; 0 (and the info log on stdout) on failure.
function gl_shader(kind: int, src: pointer) -> int {
let id = gl_create_shader(kind)
var srcs: pointers = bytes(8)
srcs[0] = src
gl_shader_source(id, 1, srcs, null)
gl_compile_shader(id)
let ids = gl_scratch()
gl_get_shaderiv(id, GL_COMPILE_STATUS, ids)
if ids[0] == 0 {
if gl_log_buf == null { gl_log_buf = bytes(8192) }
gl_get_shader_info_log(id, 8191, null, gl_log_buf)
print("shader compile failed:")
print(gl_log_buf)
gl_delete_shader(id)
return 0
}
return id
}
# Link a program from a vertex + fragment source pair; 0 on failure.
function gl_program(vs: pointer, fs: pointer) -> int {
return gl_program5(vs, null, null, null, fs)
}
# Link a program from up to five stages (null = stage absent).
function gl_program5(vs: pointer, tcs: pointer, tes: pointer, gs: pointer, fs: pointer) -> int {
let prog = gl_create_program()
var ok = true
if vs != null { let s = gl_shader(GL_VERTEX_SHADER, vs); if s == 0 { ok = false } else { gl_attach_shader(prog, s) } }
if tcs != null { let s = gl_shader(GL_TESS_CONTROL_SHADER, tcs); if s == 0 { ok = false } else { gl_attach_shader(prog, s) } }
if tes != null { let s = gl_shader(GL_TESS_EVALUATION_SHADER, tes); if s == 0 { ok = false } else { gl_attach_shader(prog, s) } }
if gs != null { let s = gl_shader(GL_GEOMETRY_SHADER, gs); if s == 0 { ok = false } else { gl_attach_shader(prog, s) } }
if fs != null { let s = gl_shader(GL_FRAGMENT_SHADER, fs); if s == 0 { ok = false } else { gl_attach_shader(prog, s) } }
if not ok { gl_delete_program(prog); return 0 }
gl_link_program(prog)
let ids = gl_scratch()
gl_get_programiv(prog, GL_LINK_STATUS, ids)
if ids[0] == 0 {
if gl_log_buf == null { gl_log_buf = bytes(8192) }
gl_get_program_info_log(prog, 8191, null, gl_log_buf)
print("program link failed:")
print(gl_log_buf)
gl_delete_program(prog)
return 0
}
return prog
}
function gl_uniform(prog: int, name: pointer) -> int { return gl_get_uniform_location(prog, name) }
# ---- buffers of floats ------------------------------------------------------------
# A float buffer is plain memory: n IEEE floats, filled from fixed (Gl.put) or
# from float bits (Gl.put_bits), uploaded with Gl.buffer_data(…, Gl.bytes_of(n), buf, …).
function gl_floats(n: int) -> pointer { return bytes(n * 4) }
function gl_bytes_of(n: int) -> int { return n * 4 }
function gl_put(buf: pointer, i: int, v: fixed) -> void { mem_put_f32(buf, i, v) }
function gl_get(buf: pointer, i: int) -> fixed { return mem_get_f32(buf, i) }
function gl_put_bits(buf: pointer, i: int, bits: int) -> void { mem_put_f32_bits(buf, i, bits) }
function gl_get_bits(buf: pointer, i: int) -> int { return mem_get_f32_bits(buf, i) }
function gl_ptr(buf: pointer, byte_offset: int) -> pointer { return mem_off(buf, byte_offset) }
function gl_f32(v: fixed) -> int { return fx_to_f32(v) }
function gl_fixed(bits: int) -> fixed { return f32_to_fx(bits) }
# One VAO, one VBO helper: create a vertex array object and return it, bound.
function gl_vao() -> int {
let ids = gl_scratch()
gl_gen_vertex_arrays(1, ids)
gl_bind_vertex_array(ids[0])
return ids[0]
}
function gl_buffer() -> int {
let ids = gl_scratch()
gl_gen_buffers(1, ids)
return ids[0]
}
function gl_texture() -> int {
let ids = gl_scratch()
gl_gen_textures(1, ids)
return ids[0]
}
function gl_framebuffer() -> int {
let ids = gl_scratch()
gl_gen_framebuffers(1, ids)
return ids[0]
}

1387
runtime/native/gl_api.ludic Normal file

File diff suppressed because it is too large Load diff

3163
runtime/native/gl_thunks.ll Normal file

File diff suppressed because it is too large Load diff

View file

@ -7,9 +7,12 @@
# a dependency to read a sprite is a poor trade when the algorithm is this
# small. So it lives here, in the language.
#
# The decoder is the canonical-Huffman formulation from Mark Adler's `puff`:
# a symbol table plus per-length counts, walked one bit at a time. Slower than
# a lookup-table decoder, and entirely fast enough to load sprites at startup.
# The decoder is the canonical-Huffman formulation from Mark Adler's `puff`: a
# symbol table plus per-length counts. Short codes (<= Z_FAST bits, which is the
# overwhelming majority) resolve in a single lookup out of a 512-entry table built
# with the symbol table; longer ones fall back to puff's walk, one bit at a time.
# The walk alone was fine for sprites, but a PBR scene inflates hundreds of
# megabytes of texture at load, and there the table is worth its 2 KB.
# ============================================================================
# ---- bit reader (DEFLATE packs bits least-significant-first) ---------------
@ -21,6 +24,7 @@ var z_bitcnt: int = 0
var z_err: int = 0
function z_start(src: pointer, len: int) -> void {
z_tables_once()
z_src = src
z_len = len
z_pos = 0
@ -29,6 +33,17 @@ function z_start(src: pointer, len: int) -> void {
z_err = 0
}
# Fill the bit buffer to at least `n` bits without consuming any (n <= 16, so the
# buffer never shifts a byte past bit 15 and cannot reach the sign bit).
function z_need(n: int) -> void {
while z_bitcnt < n {
if z_pos >= z_len { return }
z_bitbuf = (z_bitbuf | (z_src[z_pos] << z_bitcnt))
z_pos += 1
z_bitcnt += 8
}
}
function z_bits(need: int) -> int {
var val = z_bitbuf
while z_bitcnt < need {
@ -46,10 +61,17 @@ function z_bits(need: int) -> int {
}
# ---- Huffman tables -------------------------------------------------------
# A table is a single buffer: 16 length-counts followed by the symbols in
# canonical order. One allocation, no structs.
# One buffer per table: 16 length-counts, then a Z_FASTSZ-entry lookup keyed by
# the next Z_FAST bits of the stream, then the symbols in canonical order.
# A lookup entry is (length << 16) | symbol, or 0 when no code that short matches.
# Z_FAST = 10 measured fastest over a 493 MB corpus (9 and 11 are both ~8% slower:
# 9 misses the table more often, 11 spends more clearing it per dynamic block).
const Z_FAST: int = 10
const Z_FASTSZ: int = 1024 # 1 << Z_FAST
const Z_SYMS: int = 1040 # 16 + Z_FASTSZ: where the symbols start
function z_table_new(nsym: int) -> pointer {
return words((16 + nsym))
return words((Z_SYMS + nsym))
}
# lengths[i] = code length of symbol i (0 = symbol unused)
@ -71,14 +93,65 @@ function z_table_build(table: words, lengths: words, n: int) -> void {
for s in 0 .. n {
let l = lengths[s]
if l != 0 {
table[16 + offs[l]] = s
table[Z_SYMS + offs[l]] = s
offs[l] += 1
}
}
free(offs)
# ---- the fast lookup ----
for i in 0 .. Z_FASTSZ {
table[16 + i] = 0
}
# first canonical code of each length
let firstc: words = words(17)
var code = 0
for l in 1 .. 16 {
code = ((code + table[l - 1]) << 1)
firstc[l] = code
}
var idx = 0
for l in 1 .. 16 {
let cnt = table[l]
var k = 0
while k < cnt {
let sym = table[Z_SYMS + idx]
if l <= Z_FAST {
# DEFLATE reads a code most-significant-bit first out of a stream packed
# least-significant-bit first, so the table is keyed by the reversed code
let c = firstc[l] + k
var rev = 0
var b = 0
while b < l {
rev = ((rev << 1) | ((c >> b) & 1))
b += 1
}
let entry = ((l << 16) | sym)
var j = rev
while j < Z_FASTSZ {
table[16 + j] = entry
j += (1 << l)
}
}
idx += 1
k += 1
}
}
free(firstc)
}
function z_decode(table: words) -> int {
z_need(Z_FAST)
if z_bitcnt >= Z_FAST {
let e = table[16 + (z_bitbuf & (Z_FASTSZ - 1))]
if e != 0 {
let l = (e >> 16)
z_bitbuf = (z_bitbuf >> l)
z_bitcnt -= l
return (e & 65535)
}
}
# a code longer than Z_FAST bits (or a stream too short to peek): walk it
var code = 0
var first = 0
var index = 0
@ -86,7 +159,7 @@ function z_decode(table: words) -> int {
code = (code | z_bits(1))
let count = table[len]
if code - first < count {
return table[16 + index + (code - first)]
return table[Z_SYMS + index + (code - first)]
}
index += count
first = ((first + count) << 1)
@ -123,6 +196,29 @@ function z_dist_extra(sym: int) -> int {
return (sym - 2) / 2
}
# The RFC tables above are pure functions of the symbol; compute them once rather
# than dividing per match.
var z_lbase: words = null
var z_lext: words = null
var z_dbase: words = null
var z_dext: words = null
function z_tables_once() -> void {
if z_lbase != null { return }
z_lbase = words(29)
z_lext = words(29)
for s in 0 .. 29 {
z_lbase[s] = z_len_base(s)
z_lext[s] = z_len_extra(s)
}
z_dbase = words(30)
z_dext = words(30)
for s in 0 .. 30 {
z_dbase[s] = z_dist_base(s)
z_dext[s] = z_dist_extra(s)
}
}
# ---- block decoders -------------------------------------------------------
# `out` is the destination window; returns the new write position, or -1.
function z_stored(out: pointer, at: int, cap: int) -> int {
@ -156,15 +252,20 @@ function z_codes(out: pointer, at: int, cap: int, lit: pointer, dist: pointer) -
if sym > 256 {
let s = sym - 257
if s >= 29 { return -1 }
let length = z_len_base(s) + z_bits(z_len_extra(s))
let length = z_lbase[s] + z_bits(z_lext[s])
let d = z_decode(dist)
if d < 0 { return -1 }
let distance = z_dist_base(d) + z_bits(z_dist_extra(d))
if d >= 30 { return -1 }
let distance = z_dbase[d] + z_bits(z_dext[d])
if distance > w { return -1 }
for k in 0 .. length {
if w >= cap { return -1 }
out[w] = out[w - distance]
if w + length > cap { return -1 } # bounds once, not per byte
var sp = w - distance
var k = 0
while k < length {
out[w] = out[sp]
w += 1
sp += 1
k += 1
}
}
sym = z_decode(lit)

View file

@ -223,6 +223,9 @@ var in_mx0: int = 0 # x/y at the previous frame (for the delta)
var in_my0: int = 0
var in_mdx: int = 0 # delta this frame
var in_mdy: int = 0
var in_rdx: int = 0 # the raw motion the platform reports while captured
var in_rdy: int = 0
var in_cursor_mode: int = 0
var in_mbtn: int = 0 # button bitmask (bit 0 left, 1 right, 2 middle)
var in_wheel: int = 0 # wheel delta this frame
# gamepads: connected flag, button bitmask, and IN_AXES fixed axes each
@ -298,9 +301,11 @@ function input_device_commit(k: int, replaying: int) -> void {
# headless, from the single polled key. Injection (in_sim) is OR-ed on top.
if is_windowed() {
win_held(in_dev)
let mbuf = words(4) # [x, y, button-mask, wheel]
let mbuf = words(6) # [x, y, button-mask, wheel, raw dx, raw dy]
mbuf[4] = 0; mbuf[5] = 0
win_mouse(mbuf)
in_mx = mbuf[0]; in_my = mbuf[1]; in_mbtn = mbuf[2]; in_wheel = mbuf[3]
in_rdx = mbuf[4]; in_rdy = mbuf[5]
# #51 — feed the platform gamepad + touch state into the same buffers the
# read APIs use. Each is windowed-only glue (win_pad / win_touch are DCE'd
# in a headless build); on hardware they overwrite the injected state.
@ -334,6 +339,8 @@ function input_device_commit(k: int, replaying: int) -> void {
# platform above when windowed, by Input.set_mouse before this poll otherwise).
in_mdx = in_mx - in_mx0
in_mdy = in_my - in_my0
# captured (mode 2): the cursor is a clamped reticle, the motion is the raw delta
if is_windowed() and in_cursor_mode == 2 { in_mdx = in_rdx; in_mdy = in_rdy }
in_mx0 = in_mx
in_my0 = in_my
}
@ -438,6 +445,7 @@ enum CursorMode { Normal, Hidden, Locked, Confined } # Input.cursor_mode(mode:
enum PadButton { A, B, X, Y, LeftShoulder, RightShoulder, Back, Start } # Input.bind_pad(button:) / pad_button
enum MouseButton { Left, Right, Middle } # Input.mouse_down(button:)
function input_cursor_mode(mode: int) -> void {
in_cursor_mode = mode
if is_windowed() { win_cursor_mode(mode) }
}

View file

@ -221,12 +221,40 @@ function jp_number(p: JP) -> Val {
di -= 1
}
var raw = ip * 65536 + frac
raw = jp_exponent(p, raw)
if neg != 0 { raw = -raw }
return value_fixed(raw)
}
if p.i < p.n and (p.s[p.i] == 'e' or p.s[p.i] == 'E') { # 1e-05: an exponent makes it a fixed
var raw = jp_exponent(p, ip * 65536)
if neg != 0 { raw = -raw }
return value_fixed(raw)
}
if neg != 0 { ip = -ip }
return value_int(ip)
}
# an optional exponent after a number's digits, applied to a raw Q16.16 value. Exporters
# write noise like 7.49e-09 for a zero; a fixed rounds that to 0, which is what it was.
function jp_exponent(p: JP, raw0: int) -> int {
var raw = raw0
if p.i >= p.n or (p.s[p.i] != 'e' and p.s[p.i] != 'E') { return raw }
p.i += 1
var eneg = 0
if p.i < p.n and p.s[p.i] == '-' { eneg = 1; p.i += 1 }
else if p.i < p.n and p.s[p.i] == '+' { p.i += 1 }
var e = 0
while p.i < p.n and p.s[p.i] >= '0' and p.s[p.i] <= '9' {
e = e * 10 + (p.s[p.i] - 48)
p.i += 1
}
if e > 12 { e = 12 }
var k = 0
while k < e {
if eneg != 0 { raw = raw / 10 } else { raw = raw * 10 }
k += 1
}
return raw
}
function jp_list(p: JP) -> Val {
let out = value_list()

View file

@ -875,6 +875,20 @@ function emit_ns_call(ns: pointer, meth: pointer, e: Node) -> Val {
# #62: a package-provided namespace (declared with @Namespace(Foo)) that none
# of the hardcoded core blocks matched — alias Foo.method to the bare function
# foo_method (positional args), the same generic path the core aliases use.
# Gl.* — OpenGL (runtime/native/gl.ludic + the generated gl_api.ludic): the
# 478 gl3.h entry points bound as externs gl_<snake_name>, plus the Ludic
# helpers (gl_open / gl_swap / gl_screenshot / gl_program / …). Labels are the
# declaration's parameter names, so `Gl.clear_color(red: 0.1, …)` works.
if (ns == "Gl") and (bare == null) {
bare = "gl_" + meth
var gdecl = find_fn(bare)
if (gdecl == null) { gdecl = find_extern(bare) }
if (gdecl != null) {
let gnames = param_labels(gdecl)
var gi = 0
while gi < len(gnames) { push(labels, gnames[gi]); gi += 1 }
}
}
if (bare == null) and is_registered_namespace(ns) {
bare = ns_lower(ns) + ("_") + meth
# #76: a namespace declared with a block controls its public surface — an

View file

@ -249,6 +249,8 @@ function p_postfix() -> Node {
if e.a.kind == E_ID and e.a.s == "Audio" { g_uses_audio = true }
# Http.* (#6) — any Http method splices the HTTP client runtime.
if e.a.kind == E_ID and e.a.s == "Http" { g_uses_http = true }
# Gl.* — any Gl method splices the OpenGL runtime (and links the GL backend).
if e.a.kind == E_ID and e.a.s == "Gl" { g_uses_gl = true }
# Tween.to/chain/delay/value/stop/parallel (#48): the fluent stateful handles
# live in tween.ludic, advanced by an engine-owned system each Update tick.
if e.a.kind == E_ID and e.a.s == "Tween" and (e.s == "to" or e.s == "chain" or e.s == "delay" or e.s == "value" or e.s == "stop" or e.s == "parallel") { g_uses_tween_rt = true }
@ -636,6 +638,7 @@ var g_uses_tween_rt: bool = false # Tween.to/chain/delay/… (#48) -> splice tw
var g_uses_fx: bool = false # Fx.sparks/number/clear -> splice fx.ludic; fx_tick each Update, fx_draw each Render
var g_uses_audio: bool = false # Audio.* (#22) -> splice audio.ludic; a windowed build also links audio.ll + AVFoundation
var g_uses_http: bool = false # Http.* (#6) -> splice http.ludic; links http.ll + Foundation (macOS)
var g_uses_gl: bool = false # Gl.* -> splice gl.ludic (+ generated gl_api.ludic); links gl.ll + gl_thunks.ll + OpenGL
# issue #64: functions marked @System(Phase) in a prebuilt binary module — the
# compiler registers each with the host at load (it supplies the fn address,
# which Ludic source cannot take). Parallel arrays: fn name -> phase name.
@ -1123,6 +1126,14 @@ function maybe_splice_runtime() -> void {
do_import("runtime/native/http.ludic")
cur_dir = saved
}
# Gl.*: splice the OpenGL surface (gl.ludic + the generated gl_api.ludic). The
# native calls are the linked GL entry points themselves; the window attach is
# is_windowed()-guarded, so a headless build renders into an offscreen context.
if g_uses_gl {
cur_dir = ""
do_import("runtime/native/gl.ludic")
cur_dir = saved
}
# Anim.play/Motion.to sugar (#48): the writes live in systems.ludic and use the
# reflection ABI, so splice it and force the world table even when the game does
# not otherwise trip uses_engine_systems.
@ -1193,6 +1204,7 @@ function parse_program() -> void {
g_onlisten = new []Node
g_toggled_layers = new []pointer
g_uses_regex = false
g_uses_gl = false
g_uses_bignum = false
g_uses_dict = false
g_uses_numeric = false

File diff suppressed because it is too large Load diff

View file

@ -180,6 +180,13 @@ entry {
let http = join_path(home, "runtime/native/http.ll")
cmd = `{cmd} {http} -Wl,-needed_framework,Foundation`
}
# Gl.* links the OpenGL backend: gl.ll (offscreen contexts, float helpers), the
# generated per-entry-point ABI thunks, and OpenGL.framework. Windowed or not.
if g_uses_gl {
let gl = join_path(home, "runtime/native/gl.ll")
let glt = join_path(home, "runtime/native/gl_thunks.ll")
cmd = `{cmd} {gl} {glt} -framework OpenGL`
}
cmd = `{cmd} -o {out}`
let rc = run(cmd)

View file

@ -0,0 +1,73 @@
# assets.ludic — fetch the CC0 Poly Haven assets the renderer uses
# (ported off tools/glgen/fetch_assets.sh, the last shell script in the tree).
#
# ludic assets [--force] (a game, fetching what the renderer needs)
# ludic dev fetch-assets [--force] (the same thing, from a Ludic checkout)
#
# The manifest is `packages/ludic.render3d/assets.manifest` — one `<relative path> <url>`
# pair per line — and it belongs to the renderer, not to any one project: the renderer
# decides which scanned materials it wants. It ships with the package (and so with the
# toolchain), so a game outside this repo fetches the right set without keeping its own
# copy of the list. The files land in the PROJECT, under assets/polyhaven/: they are
# large and redistributable from their origin, so they are fetched rather than tracked
# by anyone.
# the manifest that belongs to the renderer, in a checkout or in the install
function assets_manifest_path() -> pointer {
let local = "packages/ludic.render3d/assets.manifest"
if file_exists(local) { return local }
var home = getenv("LUDIC_HOME")
if home == null or slen(home) == 0 { home = `{getenv_or("HOME", "")}/.ludic` }
let inst = `{home}/packages/ludic.render3d/assets.manifest`
if file_exists(inst) { return inst }
return ""
}
function cmd_fetch_assets() -> int {
let force = arg_count() > 2 and arg(2) == "--force"
let root = "assets/polyhaven"
let mpath = assets_manifest_path()
if mpath == "" { err("ludic assets: cannot find packages/ludic.render3d/assets.manifest (is the toolchain installed?)\n"); return 1 }
let manifest = read_file(mpath)
if manifest == null { err(`ludic assets: cannot read {mpath}\n`); return 1 }
var got = 0
var had = 0
var failed = 0
let n = slen(manifest)
var i = 0
while i < n {
var j = i
while j < n and manifest[j] != 10 { j += 1 }
let line = Text.trim(str_sub(manifest, i, j))
i = j + 1
if slen(line) == 0 { continue }
# split the line into <relative path> <url> on the first run of whitespace
let ln = slen(line)
var k = 0
while k < ln and not str_space(line[k]) { k += 1 }
let rel = str_sub(line, 0, k)
while k < ln and str_space(line[k]) { k += 1 }
let url = Text.trim(str_sub(line, k, ln))
if slen(rel) == 0 or slen(url) == 0 { continue }
let dest = `{root}/{rel}`
if not force and shq(`test -s {dest}`) {
had += 1
continue
}
print(`fetch {rel}`)
run(`mkdir -p "$(dirname {dest})"`)
if not shq(`curl -sSL --retry 3 -o {dest} {url}`) {
err(`fetch-assets: cannot fetch {rel}\n`)
run(`rm -f {dest}`)
failed += 1
} else {
got += 1
}
}
if failed > 0 { err(`fetch-assets: {string(failed)} failed\n`); return 1 }
print(`OK {string(got)} fetched, {string(had)} already present`)
return 0
}

View file

@ -36,7 +36,7 @@ function compile_app(src: pointer, out: pointer, mode: int, save: bool) -> bool
if mode == 2 {
if not shq(`{ludicc()} --headless {src} --emit-llvm -o {ll}`) { return false }
if not shq(`{cc()} -O2 {ll}{pbf} -o {out}`) { return false }
if not shq(`{cc()} -O2 {ll}{gl_link_flags(ll)}{pbf} -o {out}`) { return false }
if not save { run(`rm -f {ll}`) }
return true
}
@ -46,11 +46,19 @@ function compile_app(src: pointer, out: pointer, mode: int, save: bool) -> bool
# canonical `ludicc -o` path links it only when Audio.* is used.
let cocoa = `{home}runtime/native/cocoa.ll`
let audio = `{home}runtime/native/audio.ll`
if not shq(`{cc()} -O2 {ll} {cocoa} {audio} -framework Cocoa -Wl,-needed_framework,GameController -Wl,-needed_framework,AVFoundation -Wl,-rpath,@loader_path{pbf} -o {out}`) { return false }
if not shq(`{cc()} -O2 {ll} {cocoa} {audio} -framework Cocoa -Wl,-needed_framework,GameController -Wl,-needed_framework,AVFoundation -Wl,-rpath,@loader_path{gl_link_flags(ll)}{pbf} -o {out}`) { return false }
if not save { run(`rm -f {ll}`) }
return true
}
# A program that uses Gl.* references the @lgl_* thunks; link the OpenGL backend
# (gl.ll + gl_thunks.ll + OpenGL.framework) only then, so other builds are untouched.
function gl_link_flags(ll: pointer) -> pointer {
if not shq(`grep -q "@lgl_" {ll}`) { return "" }
let home = ludic_home()
return ` {home}runtime/native/gl.ll {home}runtime/native/gl_thunks.ll -framework OpenGL`
}
# the directory part of a path, without the trailing '/' ("" when there is none)
function dir_of_path(p: pointer) -> pointer {
var last = -1

View file

@ -32,6 +32,8 @@ program LudicDev {
import "docgen.ludic"
import "docgen_gen.ludic"
import "docgen_check.ludic"
import "glgen.ludic"
import "assets.ludic"
import "release.ludic"
import "pkg.ludic"
import "pkg_test.ludic"
@ -47,6 +49,7 @@ program LudicDev {
print(" build-cli build just bin/ludicc from the IR seed")
print(" tools [--install] [--test] build the editor toolchain (ludic-fmt, ludic-lsp)")
print(" clean remove build/")
print(" fetch-assets [--force] fetch the CC0 Poly Haven assets the rendering examples use")
print("")
print("test:")
print(" test the full regression suite")
@ -65,6 +68,7 @@ program LudicDev {
print(" docs-gen [--out DIR] generate the documentation site (default build/pages)")
print(" docs-check [DIR] coverage/integrity guard over a generated docs site")
print(" docs-palette [--check] regenerate emit_color.ludic + palette.json from the palette table")
print(" glgen [--check] regenerate gl_api.ludic + gl_thunks.ll from the platform gl3.h")
print("")
print("release:")
print(" release [major|minor|patch] [--dry-run] [--publish]")
@ -101,6 +105,8 @@ program LudicDev {
if (cmd == "check-vocabulary") { return cmd_check_vocab() }
if (cmd == "lint-asset") { return cmd_lint_asset() }
if (cmd == "docs-palette") { return cmd_docs_palette() }
if (cmd == "glgen") { return cmd_glgen() }
if (cmd == "fetch-assets") { return cmd_fetch_assets() }
if (cmd == "docs-gen") { return cmd_docs_gen() }
if (cmd == "docs-check") { return cmd_docs_check() }
if (cmd == "golden") { return cmd_golden() }

542
tools/ludic-cli/glgen.ludic Normal file
View file

@ -0,0 +1,542 @@
# glgen.ludic — the OpenGL binding generator, in Ludic (ported off glgen.py, the
# same way docgen.ludic was ported off gen.py). One x subcommand:
#
# ludic-dev glgen [--check] read the platform gl3.h and emit
# runtime/native/gl_api.ludic + runtime/native/gl_thunks.ll;
# --check regenerates into scratch files and compares,
# so the drift guard judges the working tree.
#
# Both outputs are tracked, so a build never runs this: it is the tool you run
# when the SDK's gl3.h changes. Output is byte-identical to the Python generator
# it replaces (verified by the drift guard in `ludic-dev test`).
# ---- parsed constants -------------------------------------------------------
var glg_cname: []pointer = null # GL_DEPTH_BUFFER_BIT
var glg_ctype: []pointer = null # "int" | "long"
var glg_cval: []pointer = null # the literal as it is emitted
# ---- parsed entry points ----------------------------------------------------
# Parallel arrays; the params of function f are the pcount[f] entries of the flat
# pp_* arrays starting at poff[f]. Flat arrays keep this free of nested slices.
var glg_fname: []pointer = null # "CullFace" (the gl prefix already stripped)
var glg_frbase: []pointer = null # return base type ("void", "GLuint", …)
var glg_frstar: []int = null # return pointer depth
var glg_fpoff: []int = null
var glg_fpcnt: []int = null
var glg_ppbase: []pointer = null
var glg_ppstar: []int = null
var glg_ppname: []pointer = null
# ---- character helpers ------------------------------------------------------
function glg_upper(c: int) -> bool { return c >= 'A' and c <= 'Z' }
function glg_lower(c: int) -> bool { return c >= 'a' and c <= 'z' }
function glg_digit(c: int) -> bool { return c >= '0' and c <= '9' }
function glg_hexdig(c: int) -> bool {
return glg_digit(c) or (c >= 'a' and c <= 'f') or (c >= 'A' and c <= 'F')
}
# a GL_* macro name character
function glg_namech(c: int) -> bool { return glg_upper(c) or glg_digit(c) or c == '_' }
# a fresh NUL-terminated copy of s[a..b)
# CamelCase -> snake_case, exactly as glgen.py's snake(): an underscore goes in
# before an upper-case letter that follows a lower-case one, nowhere else.
function glg_snake(n: pointer) -> pointer {
let b = sb_new()
let m = slen(n)
var i = 0
while i < m {
let c = n[i]
if glg_upper(c) and i > 0 and glg_lower(n[i - 1]) { sb_putc(b, '_') }
if glg_upper(c) { sb_putc(b, c + 32) } else { sb_putc(b, c) }
i += 1
}
return sb_str(b)
}
# ---- the type tables (glgen.py's INT / LONG / FLT / DBL) --------------------
# Membership in a space-delimited list, matched as whole tokens — one line per
# table instead of a chain of `or`s that a newline would cut in half.
function glg_in(t: pointer, set: pointer) -> bool {
let n = slen(set)
var p = 0
while p <= n {
var q = p
while q < n and set[q] != 32 { q += 1 }
if str_sub(set, p, q) == t { return true }
if q >= n { break }
p = q + 1
}
return false
}
function glg_is_int(t: pointer) -> bool {
return glg_in(t, "GLenum GLuint GLint GLsizei GLbitfield GLshort GLushort GLbyte GLubyte GLhalf GLchar GLboolean GLfixed")
}
function glg_is_long(t: pointer) -> bool {
return glg_in(t, "GLsizeiptr GLintptr GLint64 GLuint64 GLint64EXT GLuint64EXT")
}
function glg_is_flt(t: pointer) -> bool { return glg_in(t, "GLfloat GLclampf") }
function glg_is_dbl(t: pointer) -> bool { return glg_in(t, "GLdouble GLclampd") }
# the LLVM type a C parameter of this shape has
function glg_ir_ty(base: pointer, star: int) -> pointer {
if star > 0 { return "ptr" }
if base == "GLboolean" { return "i8" }
if base == "GLbyte" or base == "GLubyte" or base == "GLchar" { return "i8" }
if base == "GLshort" or base == "GLushort" or base == "GLhalf" { return "i16" }
if glg_is_int(base) { return "i32" }
if glg_is_long(base) { return "i64" }
if glg_is_flt(base) { return "float" }
if glg_is_dbl(base) { return "double" }
if base == "GLsync" { return "ptr" }
if base == "void" { return "void" }
return null
}
# the Ludic type it is bound as
function glg_ludic_ty(base: pointer, star: int) -> pointer {
if star > 0 or base == "GLsync" { return "pointer" }
if glg_is_int(base) { return "int" }
if glg_is_long(base) { return "long" }
if glg_is_flt(base) or glg_is_dbl(base) { return "fixed" }
if base == "void" { return "void" }
return null
}
# the Ludic ABI type carrying an LLVM type across the extern boundary
function glg_abi_ty(ct: pointer) -> pointer {
if ct == "i64" { return "i64" }
if ct == "ptr" { return "ptr" }
if ct == "void" { return "void" }
return "i32"
}
# ---- parameter names --------------------------------------------------------
# A Ludic keyword or type name cannot label a parameter. glVertexAttribPointer's
# `pointer` is a byte offset, so it says so rather than wearing a trailing _.
function glg_reserved(n: pointer) -> bool {
if glg_in(n, "program import property model enum ui namespace const var function extern handler entry") { return true }
if glg_in(n, "event scene test phase query on cancellable public layer start") { return true }
if glg_in(n, "let return if else while for in spawn despawn enable disable match machine state become where prefab") { return true }
if glg_in(n, "and or not break continue new emit cancel try") { return true }
if glg_in(n, "int long fixed countdown bool entity string pointer byte words fixeds pointers") { return true }
if glg_in(n, "Vector IVec2 Rect void true false null") { return true }
return false
}
function glg_pname(raw: pointer, i: int) -> pointer {
var n = raw
if slen(n) == 0 { n = `a{string(i)}` }
n = glg_snake(n)
if n == "pointer" { return "offset" }
if n == "program" { return "prog" }
if n == "start" { return "first" }
if n == "string" { return "text" }
if n == "layer" { return "level" }
if glg_reserved(n) { return n + "_" }
return n
}
# ---- one C declarator -> (base, star, name) ---------------------------------
# glgen.py's ctype_of: count the stars, drop every `const`, split what is left.
# Sets glg_t_base / glg_t_star / glg_t_name (Ludic has no tuple return).
var glg_t_base: pointer = null
var glg_t_star: int = 0
var glg_t_name: pointer = null
function glg_ctype_of(decl: pointer) -> void {
let m = slen(decl)
var star = 0
var i = 0
while i < m {
if decl[i] == '*' { star += 1 }
i += 1
}
# drop `const` (whole words) and turn '*' into a separator
let b = sb_new()
i = 0
while i < m {
if decl[i] == 'c' and i + 5 <= m and str_sub(decl, i, i + 5) == "const" {
i += 5
sb_putc(b, 32)
} else {
if decl[i] == '*' { sb_putc(b, 32) } else { sb_putc(b, decl[i]) }
i += 1
}
}
let flat = sb_str(b)
# split on whitespace
let parts = new []pointer
let fn = slen(flat)
var p = 0
while p < fn {
while p < fn and str_space(flat[p]) { p += 1 }
if p >= fn { break }
var q = p
while q < fn and not str_space(flat[q]) { q += 1 }
push(parts, str_sub(flat, p, q))
p = q
}
glg_t_star = star
if len(parts) == 0 { glg_t_base = ""; glg_t_name = ""; return }
glg_t_base = parts[0]
if len(parts) > 1 { glg_t_name = parts[1] } else { glg_t_name = "" }
}
# ---- constants --------------------------------------------------------------
# `#define GL_NAME <0xHEX|-?DEC>[u|U|ull|ULL]` and nothing else on the line —
# the shape glgen.py's regex accepted. Returns false when the line is not one.
var glg_v_hex: bool = false
var glg_v_digits: pointer = null # the numeric text, suffix stripped
var glg_v_neg: bool = false
function glg_num_of(tok: pointer) -> bool {
let n = slen(tok)
var i = 0
glg_v_neg = false
glg_v_hex = false
if i < n and tok[i] == '-' { glg_v_neg = true; i += 1 }
let numstart = i
if i + 1 < n and tok[i] == '0' and tok[i + 1] == 'x' {
if glg_v_neg { return false }
glg_v_hex = true
i += 2
let ds = i
while i < n and glg_hexdig(tok[i]) { i += 1 }
if i == ds { return false }
} else {
let ds = i
while i < n and glg_digit(tok[i]) { i += 1 }
if i == ds { return false }
}
glg_v_digits = str_sub(tok, numstart, i)
# what is left must be one of the accepted integer suffixes
let suf = str_sub(tok, i, n)
if suf == "" or suf == "u" or suf == "U" or suf == "ull" or suf == "ULL" { return true }
return false
}
# accumulate hex digits into an i32: the wrap makes an 8-digit value with the top
# bit set come out as the signed integer with those exact bits (glgen.py's v - 2^32).
function glg_hex_i32(h: pointer) -> int {
var v = 0
var i = 0
while i < slen(h) {
let c = h[i]
var d = 0
if glg_digit(c) { d = c - '0' }
if c >= 'a' and c <= 'f' { d = c - 'a' + 10 }
if c >= 'A' and c <= 'F' { d = c - 'A' + 10 }
v = v * 16 + d
i += 1
}
return v
}
function glg_parse_define(line: pointer) -> void {
let n = slen(line)
if n < 8 { return }
if str_sub(line, 0, 8) != "#define " { return }
var i = 8
while i < n and str_space(line[i]) { i += 1 }
let ns = i
while i < n and glg_namech(line[i]) { i += 1 }
let name = str_sub(line, ns, i)
if slen(name) < 4 { return }
if str_sub(name, 0, 3) != "GL_" { return }
# the name must be followed by whitespace (not '(' — that is a macro)
if i >= n or not str_space(line[i]) { return }
while i < n and str_space(line[i]) { i += 1 }
let vs = i
while i < n and not str_space(line[i]) { i += 1 }
let tok = str_sub(line, vs, i)
# trailing whitespace only
while i < n and str_space(line[i]) { i += 1 }
if i != n { return }
if slen(tok) == 0 { return }
# glgen.py skips the GL_VERSION_* feature macros and any repeated name
if slen(name) >= 11 and str_sub(name, 0, 11) == "GL_VERSION_" { return }
var k = 0
while k < len(glg_cname) {
if glg_cname[k] == name { return }
k += 1
}
if not glg_num_of(tok) { return }
# classify the width exactly as glgen.py did, but by digit count so a 64-bit
# all-ones literal never has to be parsed into a signed 64-bit register.
if glg_v_hex {
var d = glg_v_digits
var h = str_sub(d, 2, slen(d)) # drop the 0x
var z = 0
while z < slen(h) - 1 and h[z] == '0' { z += 1 }
h = str_sub(h, z, slen(h))
let hn = slen(h)
if hn > 8 {
push(glg_cname, name); push(glg_ctype, "long"); push(glg_cval, "-1")
return
}
if hn == 8 and not (h[0] >= '0' and h[0] <= '7') {
# 0x8… — the same bits as a negative i32; emit the signed decimal
let v = glg_hex_i32(h)
push(glg_cname, name); push(glg_ctype, "int"); push(glg_cval, string(v))
return
}
push(glg_cname, name); push(glg_ctype, "int"); push(glg_cval, glg_v_digits)
return
}
var lit = glg_v_digits
if glg_v_neg { lit = "-" + lit }
push(glg_cname, name); push(glg_ctype, "int"); push(glg_cval, lit)
}
# ---- entry points -----------------------------------------------------------
# `GLAPI <ret> APIENTRY gl<Name> (<args>) …;` — every gl3.h declaration is on one
# line, so a line scan matches what glgen.py's re.M regex did.
function glg_parse_glapi(line: pointer) -> void {
let n = slen(line)
if n < 6 { return }
if str_sub(line, 0, 6) != "GLAPI " { return }
let ap = Text.index_of(line, " APIENTRY ")
if ap < 0 { return }
let ret = Text.trim(str_sub(line, 6, ap))
var i = ap + 10
while i < n and str_space(line[i]) { i += 1 }
let ns = i
while i < n and (glg_upper(line[i]) or glg_lower(line[i]) or glg_digit(line[i]) or line[i] == '_') { i += 1 }
let fname = str_sub(line, ns, i)
if slen(fname) < 3 { return }
if str_sub(fname, 0, 2) != "gl" { return }
while i < n and str_space(line[i]) { i += 1 }
if i >= n or line[i] != '(' { return }
let argstart = i + 1
var depth = 1
i += 1
while i < n and depth > 0 {
if line[i] == '(' { depth += 1 }
if line[i] == ')' { depth -= 1 }
if depth == 0 { break }
i += 1
}
if depth != 0 { return }
let args = str_sub(line, argstart, i)
let bare = str_sub(fname, 2, slen(fname))
var k = 0
while k < len(glg_fname) {
if glg_fname[k] == bare { return }
k += 1
}
glg_ctype_of(ret)
let rbase = glg_t_base
let rstar = glg_t_star
let poff = len(glg_ppbase)
var pcnt = 0
let at = Text.trim(args)
if at != "void" and at != "" {
# split the argument list on commas (no function-pointer args in gl3.h)
let an = slen(args)
var p = 0
while p <= an {
var q = p
while q < an and args[q] != ',' { q += 1 }
let one = str_sub(args, p, q)
glg_ctype_of(one)
push(glg_ppbase, glg_t_base)
push(glg_ppstar, glg_t_star)
push(glg_ppname, glg_t_name)
pcnt += 1
if q >= an { break }
p = q + 1
}
}
push(glg_fname, bare)
push(glg_frbase, rbase)
push(glg_frstar, rstar)
push(glg_fpoff, poff)
push(glg_fpcnt, pcnt)
}
function glg_parse(h: pointer) -> void {
glg_cname = new []pointer; glg_ctype = new []pointer; glg_cval = new []pointer
glg_fname = new []pointer; glg_frbase = new []pointer; glg_frstar = new []int
glg_fpoff = new []int; glg_fpcnt = new []int
glg_ppbase = new []pointer; glg_ppstar = new []int; glg_ppname = new []pointer
let n = slen(h)
var i = 0
while i < n {
var j = i
while j < n and h[j] != 10 { j += 1 }
let line = str_sub(h, i, j)
if slen(line) > 0 {
if line[0] == '#' { glg_parse_define(line) }
if line[0] == 'G' { glg_parse_glapi(line) }
}
i = j + 1
}
}
# ---- emit gl_api.ludic ------------------------------------------------------
function glg_emit_ludic(path: pointer) -> bool {
let b = sb_new()
sb_puts(b, "# ============================================================================\n")
sb_puts(b, "# gl_api.ludic — the OpenGL 4.1 core API, bound for Ludic. GENERATED by\n")
sb_puts(b, "# `ludic-dev glgen` from the platform gl3.h: every entry point and every GL_* constant.\n")
sb_puts(b, "# Do not edit by hand; regenerate. float/double parameters take `fixed`; the IR\n")
sb_puts(b, "# thunks in runtime/native/gl.ll (also generated) convert at the C boundary.\n")
sb_puts(b, "# ============================================================================\n")
sb_puts(b, "\n")
var i = 0
while i < len(glg_cname) {
sb_puts(b, `const {glg_cname[i]}: {glg_ctype[i]} = {glg_cval[i]}\n`)
i += 1
}
sb_puts(b, "\n")
i = 0
while i < len(glg_fname) {
let lname = "gl_" + glg_snake(glg_fname[i])
let ps = sb_new()
var p = 0
while p < glg_fpcnt[i] {
let ix = glg_fpoff[i] + p
if p > 0 { sb_puts(ps, ", ") }
let lt = glg_ludic_ty(glg_ppbase[ix], glg_ppstar[ix])
sb_puts(ps, `{glg_pname(glg_ppname[ix], p)}: {lt}`)
p += 1
}
let rt = glg_ludic_ty(glg_frbase[i], glg_frstar[i])
var rets = ""
if rt != "void" { rets = " -> " + rt }
sb_puts(b, `extern function {lname}({sb_str(ps)}){rets} = "lgl_{glg_fname[i]}"\n`)
i += 1
}
return write_file(path, sb_str(b))
}
# ---- emit gl_thunks.ll ------------------------------------------------------
function glg_emit_thunks(path: pointer) -> bool {
let b = sb_new()
sb_puts(b, "; ============================================================================\n")
sb_puts(b, "; gl_thunks.ll — one thunk per OpenGL 4.1 core entry point. GENERATED by\n")
sb_puts(b, "; `ludic-dev glgen` from gl3.h. Each @lgl_* takes the Ludic ABI (i32 / i64 / ptr,\n")
sb_puts(b, "; float and double as Q16.16 fixed) and calls the real gl* with exact C types.\n")
sb_puts(b, "; ============================================================================\n")
sb_puts(b, "\n")
var i = 0
while i < len(glg_fname) {
let cret = glg_ir_ty(glg_frbase[i], glg_frstar[i])
let ps = sb_new()
var p = 0
while p < glg_fpcnt[i] {
let ix = glg_fpoff[i] + p
if p > 0 { sb_puts(ps, ", ") }
sb_puts(ps, glg_ir_ty(glg_ppbase[ix], glg_ppstar[ix]))
p += 1
}
sb_puts(b, `declare {cret} @gl{glg_fname[i]}({sb_str(ps)})\n`)
i += 1
}
sb_puts(b, "\n")
i = 0
while i < len(glg_fname) {
let cret = glg_ir_ty(glg_frbase[i], glg_frstar[i])
let lret = glg_abi_ty(cret)
let sig = sb_new()
let body = sb_new()
let call = sb_new()
var p = 0
while p < glg_fpcnt[i] {
let ix = glg_fpoff[i] + p
let ct = glg_ir_ty(glg_ppbase[ix], glg_ppstar[ix])
let lt = glg_abi_ty(ct)
if p > 0 { sb_puts(sig, ", "); sb_puts(call, ", ") }
sb_puts(sig, `{lt} %a{string(p)}`)
if ct == lt {
sb_puts(call, `{ct} %a{string(p)}`)
} else {
if ct == "i8" or ct == "i16" {
sb_puts(body, ` %c{string(p)} = trunc i32 %a{string(p)} to {ct}\n`)
sb_puts(call, `{ct} %c{string(p)}`)
} else {
# float / double: the Ludic side passes Q16.16, so scale by 1/65536
sb_puts(body, ` %f{string(p)} = sitofp i32 %a{string(p)} to {ct}\n`)
sb_puts(body, ` %c{string(p)} = fmul {ct} %f{string(p)}, 0x3EF0000000000000\n`)
sb_puts(call, `{ct} %c{string(p)}`)
}
}
p += 1
}
sb_puts(b, `define {lret} @lgl_{glg_fname[i]}({sb_str(sig)}) `)
sb_puts(b, "{\n")
sb_puts(b, "entry:\n")
sb_puts(b, sb_str(body))
let invoke = `call {cret} @gl{glg_fname[i]}({sb_str(call)})`
if cret == "void" {
sb_puts(b, ` {invoke}\n`)
sb_puts(b, " ret void\n")
} else {
if cret == lret {
sb_puts(b, ` %r = {invoke}\n`)
sb_puts(b, ` ret {lret} %r\n`)
} else {
if cret == "i8" or cret == "i16" {
sb_puts(b, ` %r = {invoke}\n`)
sb_puts(b, ` %z = zext {cret} %r to i32\n`)
sb_puts(b, " ret i32 %z\n")
} else {
sb_puts(b, ` %r = {invoke}\n`)
sb_puts(b, ` %m = fmul {cret} %r, 65536.0\n`)
sb_puts(b, ` %z = fptosi {cret} %m to i32\n`)
sb_puts(b, " ret i32 %z\n")
}
}
}
sb_puts(b, "}\n")
i += 1
}
return write_file(path, sb_str(b))
}
# ---- the task ---------------------------------------------------------------
# `ludic-dev glgen` rewrites the two tracked outputs; `--check` regenerates into
# scratch files and compares, so the guard judges the working tree, not git HEAD.
function cmd_glgen() -> int {
let check = arg_count() > 2 and arg(2) == "--check"
let sdk = Text.trim(capture("xcrun --show-sdk-path"))
if sdk == "" { err("glgen: no macOS SDK (xcrun --show-sdk-path)\n"); return 1 }
let hpath = `{sdk}/System/Library/Frameworks/OpenGL.framework/Headers/gl3.h`
let h = read_file(hpath)
if h == null { err(`glgen: cannot read {hpath}\n`); return 1 }
glg_parse(h)
if len(glg_fname) == 0 { err("glgen: no GLAPI declarations found\n"); return 1 }
# every type in the header must be one this generator knows how to bind
var i = 0
while i < len(glg_fname) {
if glg_ir_ty(glg_frbase[i], glg_frstar[i]) == null or glg_ludic_ty(glg_frbase[i], glg_frstar[i]) == null {
err(`glgen: unknown return type {glg_frbase[i]} on gl{glg_fname[i]}\n`); return 1
}
var p = 0
while p < glg_fpcnt[i] {
let ix = glg_fpoff[i] + p
if glg_ir_ty(glg_ppbase[ix], glg_ppstar[ix]) == null or glg_ludic_ty(glg_ppbase[ix], glg_ppstar[ix]) == null {
err(`glgen: unknown parameter type {glg_ppbase[ix]} on gl{glg_fname[i]}\n`); return 1
}
p += 1
}
i += 1
}
var api_out = "runtime/native/gl_api.ludic"
var thunk_out = "runtime/native/gl_thunks.ll"
if check { api_out = tmp_path("gl_api.ludic"); thunk_out = tmp_path("gl_thunks.ll") }
if not glg_emit_ludic(api_out) { err("glgen: cannot write gl_api.ludic\n"); return 1 }
if not glg_emit_thunks(thunk_out) { err("glgen: cannot write gl_thunks.ll\n"); return 1 }
if check {
if not shq(`cmp -s {api_out} runtime/native/gl_api.ludic`) { err("gl_api.ludic drifted from gl3.h (run: ludic-dev glgen)\n"); return 1 }
if not shq(`cmp -s {thunk_out} runtime/native/gl_thunks.ll`) { err("gl_thunks.ll drifted from gl3.h (run: ludic-dev glgen)\n"); return 1 }
}
print(`OK {string(len(glg_cname))} constants, {string(len(glg_fname))} entry points`)
return 0
}

View file

@ -24,6 +24,7 @@ program Ludic {
import "build.ludic"
import "project.ludic"
import "pkg.ludic"
import "assets.ludic"
function usage() -> void {
print("ludic — the toolchain for the Ludic language")
@ -44,6 +45,7 @@ program Ludic {
print(" update [module] bump a dependency (or all) to its latest published version")
print(" verify check every locked package against the store by content hash")
print(" vendor copy the resolved packages into ./vendor for offline builds")
print(" assets [--force] fetch the CC0 materials the renderer needs into assets/polyhaven/")
print(" build-lib <module.ludic> compile a package's module to a prebuilt dylib in lib/<target>/")
print(" link-flags print the clang flags to link this project's prebuilt module dylibs")
print("")
@ -85,6 +87,7 @@ program Ludic {
if (cmd == "update") { return cmd_pkg_update() }
if (cmd == "verify") { return cmd_pkg_verify() }
if (cmd == "vendor") { return cmd_pkg_vendor() }
if (cmd == "assets") { return cmd_fetch_assets() }
if (cmd == "build-lib") { return cmd_pkg_build_lib() }
if (cmd == "link-flags") { return cmd_pkg_link_flags() }

View file

@ -35,6 +35,22 @@ function write_file(path: pointer, s: pointer) -> bool {
return true
}
# ---- byte strings ------------------------------------------------------------------
# s[a, b) as a fresh NUL-terminated string; out-of-range ends are clamped.
function str_sub(s: pointer, a: int, b: int) -> pointer {
var lo = a
if lo < 0 { lo = 0 }
var hi = b
if hi < lo { hi = lo }
let out = bytes(hi - lo + 1)
var i = lo
var k = 0
while i < hi { out[k] = s[i]; k += 1; i += 1 }
out[k] = 0
return out
}
function str_space(c: int) -> bool { return c == 32 or c == 9 or c == 13 }
function file_exists(path: pointer) -> bool { return shq(`test -e {path}`) }
function is_exec(path: pointer) -> bool { return shq(`test -x {path}`) }
# is `a` newer than `b` (like the shell's `-nt`)?

View file

@ -22,14 +22,27 @@ function manifest_name() -> pointer {
# the name to give the built binary: the manifest's module (its last dotted
# segment, so ludic.snake builds `snake`), else the entry file's base name.
# The built binary's name, from the manifest's module path when there is one.
#
# A module path is a URL — `git.host/user/maroon-lake` — so the last path segment comes
# first: taking the last DOT of that would have cut inside the host and produced
# `io/user/maroon-lake`, which git-hosted names all share and which `build/{name}` then
# turned into directories. Within the segment a dot still separates a namespace from the
# package (`ludic.render3d` builds as `render3d`).
function project_name(entry: pointer) -> pointer {
let mod = manifest_name()
var mod = manifest_name()
if mod != "" {
var last = -1
var n = 0
while mod[n] != 0 { n += 1 }
var slash = -1
var i = 0
while mod[i] != 0 { if mod[i] == '.' { last = i }; i += 1 }
if last >= 0 { return mod[last + 1..i] }
return mod
while i < n { if mod[i] == '/' { slash = i }; i += 1 }
if slash >= 0 { mod = mod[slash + 1..n]; n = n - slash - 1 }
var last = -1
i = 0
while i < n { if mod[i] == '.' { last = i }; i += 1 }
if last >= 0 { return mod[last + 1..n] }
if n > 0 { return mod }
}
return capture_line(`basename {entry} .ludic`)
}

View file

@ -447,6 +447,7 @@ function cmd_dev_test() -> int {
smoke("events/events")
smoke("networking/net_rt")
smoke("library/cursor_capture") # #89 Input.cursor_mode compiles (no-op headless; windowed links cocoa.ll)
smoke("rendering/gl_triangle") # Gl.* (OpenGL 4.1 core) compiles headless; the run needs a GPU context
print("== the compiler and the CLI (ludicc / ludic) ==")
# ludicc comes out of the IR seed with clang alone; the CLI is then compiled
@ -525,5 +526,14 @@ function cmd_dev_test() -> int {
ok("ludic-dev docs-palette regenerates emit_color.ludic + palette.json byte-identically")
} else { bad2("ludic-dev docs-palette --check", capture_line(`tail -1 {tmp_dir()}/pal.out`)) }
# the OpenGL binding generator is the same shape: gl_api.ludic + gl_thunks.ll are
# tracked, and --check regenerates them from the platform gl3.h and compares. The
# header is macOS-only, so this guard cannot run off Darwin.
if is_darwin() {
if shq(`bin/ludic-dev glgen --check > {tmp_dir()}/glgen.out 2>&1`) {
ok("ludic-dev glgen regenerates gl_api.ludic + gl_thunks.ll byte-identically")
} else { bad2("ludic-dev glgen --check", capture_line(`tail -1 {tmp_dir()}/glgen.out`)) }
} else { skip("ludic-dev glgen --check (needs the macOS OpenGL headers)") }
return report()
}