Skip to content

devlog

The bin only falls one way

2026-10-03

Captured places make animated motion look fake. Capturing the motion too fixes that and introduces a harder problem.

A scanned street carries its own light. Every ellipsoid in it was measured under a real sky on a real afternoon, which is the whole reason the place reads as a photograph rather than as a model of a photograph. Then you put something in it that moves, and the thing that moves was made the ordinary way: a mesh somebody modelled, a rig somebody bound, a curve somebody drew. It does not matter how good the curve is. The scan and the animation disagree about what light is, and the eye finds the seam immediately.

The fix that interests me is not better animation. It is capturing the motion the same way the street was captured.

Four-dimensional Gaussian splatting is the ordinary kind with time in it. Instead of several million ellipsoids frozen at one instant, you get something that can be asked what the scene looked like at a given moment across a stretch of seconds. There is more than one way to build it. You can keep one canonical set of splats and learn a field that warps them as time moves, or you can let the primitives themselves carry a temporal extent and fade in and out of existence. Both end up in the same place for my purposes: the appearance and the motion come out of a single recording, so they cannot disagree with each other.

The test case I picked is a bin that gets kicked. It is small on purpose. One object, two seconds, and a trigger that already exists, because the physics is already running and already knows the instant a wheel or a foot touched it. It is also unforgiving on purpose. Everyone has seen a bin go over, so if the playback is wrong nobody needs to be told.

The hard part showed up about an hour into thinking about it, and it is not the rendering.

A captured clip is one fall. The bin I record went over because I kicked it from one side, at one speed, and it landed where it landed. In the game it can be hit from any side at any speed, and there are infinitely many ways for it to go. A library of captures does not fix that, it only narrows the gap, and the one time a player hits it from an angle I never recorded is the time the whole illusion dies.

So the direction I am taking is to stop asking the capture to do the part it is bad at. Physics is good at where an object goes: the arc, the bounce, the pose it settles into. It has always been good at that and it needs no help from a recording. What physics cannot give me is the part a capture is uniquely good at, which is how the surface behaves while all of that is happening. The dent as the metal takes the kick. The lid slapping and settling. The way light slides across a scratched panel as it turns over. Split the rigid motion from the deformation, let the simulation own the first and the capture own the second, and a clip stops being a recording you play back and becomes a description of how this particular object responds to being hit.

Whether that split survives contact with a real capture I genuinely do not know yet. It may turn out the deformation and the trajectory are entangled badly enough that pulling them apart looks worse than either on its own.

Two more problems are queued behind it, both real and both less interesting. One is sorting. A splat scene is drawn in a single global back-to-front order, and an object moving through the street has to take its place in that order every frame without the street paying for it. The other is simply what a second of captured motion costs to store and to ship, which is the question the streaming side of this project already exists to answer, and which gets harder when the thing being streamed is changing over time.

None of this is live. I am writing it down for the same reason as the last one of these: the interesting problems here are the ones that do not have a known shape yet, and a log of only the finished things describes a different job from the one I am actually doing.