A week ago I wrote that the format, the runtime and the renderer had all become ours, and I ended it by saying I would rather be argued with from the numbers than believed. Then I printed no numbers. That is a sentence that sounds like confidence and works as a way of avoiding the check, so here is the whole ledger, including the parts that lose.
The work went in four stages. Each one is measured against the path the game actually ships, on the same benches, and the rule is that the shipped path stays until the new one beats it.
Reading the data ourselves came first. The decoder is matched attribute by attribute against the engine's own reader rather than eyeballed: a real file from La Sarraz, 253,858 splats, comes out identical and takes 49 ms on one thread. One thing in it is unsolved and it is annoying rather than deep. A browser's 2D canvas premultiplies alpha, which quietly destroys the lowest opacities, so the exact decode that works outside a browser does not yet work inside one. The ways out are a GPU texture, WebCodecs, or writing a lossless WebP decoder. I have not picked.
The runtime came second: which parts of a place to hold, which to fetch next, what to show while you wait. On the shared score it beat the engine path in all seven bench scenarios, by 3 to 12 per cent at equal bytes, and by 25 per cent on a long fat link where latency dominates. Then it had to draw, and drawing produced the sharpest lesson of the month. A rotation stored in the wrong component order gave 19.9 dB and a picture that looked shredded. Stored the other way it gave 58.7 dB against the engine's own loader. The documentation said one order and the engine wanted the other, and nothing about the wrong one announces itself as a convention problem.
Drawn through the engine's own renderer, our runtime won the picture in all seven scenarios and matched the engine for smoothness on Aix. It lost on La Sarraz, and that one belongs on the record because it is the only honest way to read the rest. The slow frames there were the engine reallocating GPU buffers to exact counts as the scene changed size, which is exactly the thing a general-purpose renderer has to do and a single-purpose one does not.
So the third stage removed it. Our own pool, sized once, at 32 bytes a splat. Our own projection, our own 20-bit radix sort, our own raster. Against PlayCanvas that is 43.5 dB on a single file and about 30 dB at 8.26 million splats, the same by eye, and faster on both machines I own: 8.5 ms against 12.1 on a GTX 1050, and 56 against 74 on an Intel UHD 620. Streaming a whole town on our stack end to end gave 0 or 1 frames over 50 ms in all seven scenarios, and the best picture in six of them.
The last piece was light, and it mattered more than its size suggests, because wanting light on the splats rather than only on the meshes standing in front of them was one of the three reasons for owning the renderer at all. Against a known answer it now matches the engine within 0.3 luma everywhere.
What is not done, plainly. GPU timestamp timing, so the frame figures above are wall clock and noisier than I would like. Compositing fully into the game frame with correct depth against the car and the meshes. The browser decode path, which still keeps decoded data in ordinary memory at 40 bytes a splat. And the long tail on machines I do not own, which is the one part of this I cannot bench my way out of.
It stays behind a flag. The 30 dB at 8.26 million splats is a far-field sort ordering difference rather than a wash, and until I understand it on hardware I do not have, behind a flag is where it belongs.