The last post on this left off at the problem rather than the answer, so here is the answer I am building toward, in enough detail to be argued with.
Four-dimensional Gaussian splatting is the ordinary kind with time in it, and there are two families of it worth knowing apart. In the first you keep one canonical set of primitives and learn a deformation field that warps them as time moves, so the scene at any instant is the canonical set pushed through a function of position and time. In the second the primitives themselves carry a temporal extent, a centre in time and a lifetime, and fade in and out of existence as the clip runs. The first is compact and struggles with topology that genuinely changes. The second handles appearance and disappearance naturally and costs more to store. For a rigid object being knocked over, which deforms a little and changes topology not at all, the first is the better fit.
Either way the point is the same: appearance and motion come out of one recording, so they cannot disagree with each other the way a mesh and a scan do.
Now the problem that makes this hard, and it is not rendering. A captured clip is one fall. The bin I record went over because I kicked it from one side, at one speed, and it landed where it landed. In a game it can be hit from any side at any speed, and there are infinitely many ways for it to go. A library of captures narrows that gap and never closes it, and the one impact nobody recorded is the one that kills the illusion.
So the method is to stop asking the capture to supply the part the simulation already does well.
Concretely: solve the rigid motion out of the clip before using it. Frame by frame, fit the best rigid transform taking the primitives from their rest configuration to where they are at that instant, then express every primitive relative to that transform. What comes out is two tracks. One is a rigid trajectory, which I then throw away, because the rigid body already in the scene produces a better one that responds to how the player actually hit it. The other is the residual: what the surface did in the frame of the object itself, which is a function of being struck rather than of where the object ended up.
At playback the rigid body supplies the transform and the residual supplies the rest. The bin follows the simulation, lands where the physics says, and still dents where the metal took the kick, still has the lid slap and settle, still has light slide across a scratched panel as it turns over. One capture then covers a family of impacts instead of one, because the part that varies between impacts is simulated and the part that is expensive to fake is measured.
What I do not know yet, and it is the whole risk: whether the residual is actually separable. If the deformation depends on the precise contact rather than on the impulse, then a residual recorded from one kick will be wrong under another, and pulling the two apart may read worse than either on its own. There is a related question of how far a single residual stretches across impact directions before it stops being believable, which is an empirical number nobody can give me from an armchair.
None of this is implemented. It is written down because the shape of the problem is the interesting part and a log of only the finished things describes a different job from the one I am doing.