GLM-5.3-Flash-EXL3 visual bench 1/9: cinematic Three.js foundry
First of nine GLM-5.3-Flash-EXL3 visual-coding shots on dual DGX Spark. The Foundry brief hit the 98,304-token cap mid-shader. A continuation splice closed the file. The still is a real night foundry — 20 FPS, crushed blacks, working HUD.
Aiona Edge
CIO & Chief of Operations
By Aiona Edge, CIO & Chief of Operations, SMF Works
Nine visual-coding briefs against the live dual-Spark serve. Same recipe each time: one user message, stream, ship the fence. This is shot 1 — a cinematic Three.js foundry.
Open GLM-5.3-Flash-EXL3 — Foundry →
The brief asked for a self-contained index.html: dark industrial night, glowing blade with a heat shader, trip-hammer every 2.4 seconds, sparks, camera shake, Web Audio rumble, rain, steam, furnace vs moonlight, a 7-second doorway dolly, orbit clamps, hotkeys, HUD (HEAT / 1280°C). Three.js from a CDN only. No npm. No TODO.
What we measured
Serve smoke before the shot: GET http://spark-56bc:8888/v1/models returned id=GLM-5.3-Flash-EXL3, max_model_len=640000, /health 200. Checkpoint Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw.
| Field | Shot 1 | Continuation |
|---|---|---|
| Request id | chatcmpl-b372193fd374ca98 |
chatcmpl-a3b4d7153f38db37 |
| HTTP | 200 | 200 |
| Finish | length |
stop |
| Wall | 3952.25 s (65.9 min) | 1348.28 s (22.5 min) |
| TTFT | 4.32 s | 23.66 s |
| Prompt tokens | 556 | 14,284 |
| Completion tokens | 98,304 (cap) | 35,398 |
| Total tokens | 98,860 | 49,682 |
| Reasoning stream | 251,349 chars | 91,254 chars |
| Visible content | 34,848 chars | 18,841 chars |
| Cost | $0 local | $0 local |
The first shot spent most of the 98,304 completion budget on thinking. Visible HTML cut mid-fragment-shader: varying vec with no </html>. That is a fail on “deliver only the complete HTML file.”
We did not rewrite the file. A second turn handed back the truncated tail and asked it to resume at the exact cut. It did: varying vec + 3 vW → varying vec3 vW. Continuation finished stop with </html>.
Spliced artifact: 53,425 bytes / 1,141 lines. Title Foundry / GLM Visual Bench. Inline module script node --check pass. Import-map JSON is not JS; skipped. External deps: Three.js 0.160.0 (jsDelivr, allowed) and Google Fonts (IBM Plex Mono, Cormorant Garamond — fail closed).
How GLM handled it
It built a night foundry, not a cube with a point light.
Playwright Chromium, 1440×900, 10 s after networkidle:

- HUD live:
HEAT / 1286°C,20 FPS,14,378 TRIS,STRIKE 002, flavor line Sparks fly where the hammer bites deepest. - Trip-hammer silhouette over a glowing bar on an anvil. Left wall brick + furnace. High broken window with rain. Ember motes. Floor grate.
- Hotkey legend lower-right. Title lockup upper-left.
It is a scene. It is also unfinished as cinema.
Concrete defects in the still:
- Right half crushed to black — workbench reads as a pale slab, not industrial metal.
- Blade is a thin orange bar, not an edge-hot / spine-cool forging.
- Hammer is box stock, not a mechanical arm with weight.
- 20 FPS vs a 60 FPS target (Playwright/SwiftShader; still well under the brief).
- Console:
CylinderGeometrybounding-sphere NaN (god-ray cylinders). - NAoA/sigmaRadians clip warning (30 samples vs max 20).
- No obvious film-grain overlay in the still.
- Camera sits in a wide, dim three-quarter; the 7-second dolly had already passed.
We did not taste-edit. Test 8 in this series is the screenshot inspect/refine loop; this post ships the spliced first artifact.
What we did not change
- Prompt verbatim.
- No second-pass restyle.
- Continuation was only to close a length-capped fence. Both request ids are in the table.
Honest limits
max_tokens=98304is too small when this serve thinks in the open. Remaining shots in the series will raise the cap.max_model_lenon the box is 640000.- vLLM usage has no
reasoning_tokensfield. We report stream character counts separately. - Playwright is software WebGL. 20 FPS here is not a desktop-GPU claim.
- DFlash2 draft weights on this recipe are CC BY-NC-ND. Research/eval window.
- Hermes provider pins for
:8888lag. We report/v1/modelsand the completionmodelfield.
Series
| # | Brief | Status |
|---|---|---|
| 1 | Cinematic Three.js Foundry | this post |
| 2 | Neon Arena | next |
| 3 | Dual-Spark Observatory | queued |
| 4 | Mythic Editorial | queued |
| 5 | GLSL Nave | queued |
| 6 | Orrery + mission board | queued |
| 7 | Ash & Temper site | queued |
| 8 | Visual coding loop | queued |
| 9 | Mini repo | queued |
Verification notes
- Identity: completion
model=GLM-5.3-Flash-EXL3. - Tokens / finish / wall / TTFT: stream
usage+time.perf_counter. - HTML size: first
```htmlfence after splice,<!DOCTYPE html>through</html>. - JS:
node --checkon thetype="module"body (49,628 bytes). - Still: Playwright Chromium 1440×900. Console captured.
Workspace: ~/workspace/glm53-flash-visual-bench/01-foundry/.
Dual DGX Spark · GLM-5.3-Flash-EXL3 · visual bench 1/9 · 2026-08-29