6 October 2026 · Syamjith NK
You hand a generative video model a camera move and it is under no obligation to deliver one. The usual check is a person watching sixty clips and forming an impression. So I wrote shotdrift, which measures the camera path out of the pixels instead: a grid of tiles tracked by phase correlation, one similarity transform fitted to all of them, and pan, zoom and roll falling out of the fit frame by frame. No metadata, no depth, no solve, because none of those come with the mp4.
The headline feature was going to be the melting-background tell. Everyone who has looked at generated video knows it: geometry that invents itself between one frame and the next, a wall that reorganises while the camera passes it. A tool that fits one rigid camera to a scene has, in principle, exactly the right residual sitting there to catch it, because whatever the camera model cannot explain is the other half of the product.
That feature is not in the tool, and the reason is a number I did not expect.
The measurement has to be comparable across clips, so the comparison frame is chosen adaptively: far enough away that the amount of warp between the two frames is the same every time, which keeps the resampling residue from being a function of how fast the shot happens to move. Then you measure what is left over.
The validation material is real single-camera footage from live events, 720p and 1080p, 30p and 50p, locked-off and operated, against synthetic controls built by cropping a textured plate so the fault in each one is known. The worst control I could build for this check was a cross-dissolve between two entirely different worlds, which is the limit case of invented geometry: every pixel in the frame is simultaneously two unrelated places.
| clip | span residual |
|---|---|
| real: stage camera, a 20-frame segment | 1.062 |
| real: stage camera, a 25-frame segment | 0.867 |
| real: stage camera, operated slow zoom | 0.781 |
| real: panel with a vision mix | 0.423 |
| control: cross-dissolve between two different worlds | 0.358 |
| control: clean push-in | 0.114 |
| control: clean pan | 0.107 |
Read the order. Real footage is at the top. The thing the check exists to catch sits below all of it, closer to a clean synthetic push-in than to any real camera. The ordering is not marginal and it is not noise: it is reversed, by a factor of three.
The explanation is obvious once you see the number, and it is the reason no amount of tuning rescues this. Real scenes are full of things that are not the camera. People walk through them. LED walls behind a stage change their entire content on a cue. A vision mix cuts the picture out from under you. All of that is structural change, and there is more of it in a normal event clip than there is in geometry that melts. The melting tell is a thing your eye catches instantly and a thing this measurement cannot distinguish from Tuesday.
There is no threshold in either direction, so there is no version of this check worth shipping. It is deleted. What is in the repository instead is a calibration document that records the distribution, states the failure, and keeps the dissolve control as a test.
That last part matters more than the deletion. The dissolve under a locked-off camera comes back
clean, and a test asserts that it keeps coming back clean. It looks
backwards to pin your own blind spot as a passing test, but the alternative is worse: a known gap
that nobody wrote down slowly becomes a false positive the first time somebody tightens an unrelated
bound, and then the tool is quietly flagging dissolves for a reason nobody can reconstruct. A stated
limitation with a test behind it stays a limitation. An unstated one turns into a mystery.
The residual did not go to waste, either. Per-frame structural change is a poor melt detector and
an excellent cut detector, which is the job it now does: real footage within a single shot
reaches 0.36, and a transition that preserves some structure, a whip or a dissolve, spikes to
between 1.7 and 8.6. Across a clean hard cut between unrelated shots nothing correlates at all, no
camera can be fitted, and the value is NaN. The first version only tested for the
spike, so it caught smeared transitions and walked straight past an ordinary cut.
The second idea was reversals. If a camera move keeps changing its mind, something is wrong with it. This is the kind of metric that feels unarguable until you point it at material.
| source | reversals |
|---|---|
| real: stage camera, operated slow zoom | 27 |
| real: panel with a vision mix | 16 |
| control: a push-in that deliberately reverses twice | 3 |
An operator riding a zoom rocker reverses constantly, because that is what holding a shot on a live subject looks like from the inside. The real material scores nine times worse than the deliberate fault. Normalising by clip length makes it worse rather than better, since the long clips are the operated ones. Two checks in, and both of my instincts about what "wrong" looks like had been contradicted by the same property of real footage.
So reversals are still reported, labelled (measured, not judged), and nothing is
gated on them.
Both dead checks were asking an unanswerable question: is this video wrong? The question that is answerable is narrower and more useful. Did the move you asked for happen?
$ shotdrift pan_demo.mp4 --expect push-in
camera path
dominant move pan
pan 0.7159 of frame width
zoom 1.000x
reversals 0 (measured, not judged)
incoherence 0.0000
confidence 0.98
asked for: push-in
NOT HELD - asked for push in / dolly in; zoom moved +0.0001, under the 0.02
floor - that move did not happen
no findings: the motion is consistent with one physical camera.
That clip is a perfectly good shot. Smooth, coherent, one physical camera, nothing wrong with it at all. It is also not the shot that was ordered: it pans, and the push-in never happened. Both things are true and they are reported separately, because they are different questions and they have different fixes.
There are fourteen declared moves, and each is checked three ways, because the three failures are not the same failure: did it happen at all, was it the right way round, and was it held, or did a quarter of the travel run backwards. That last one is what the reversal count was reaching for all along, and it works here only because the declared move supplies the direction that makes "backwards" mean something. Without knowing what was asked for, a reversal is suspicious. With it, a reversal is measurable.
The signs are camera-relative, which produced the silliest bug in the project. A camera panning
right makes the picture move left, so pan-right expects a negative picture-x. I wrote
those the intuitive way round, and a correctly measured, perfectly clean pan reported "x went the
other way". The measurement was right and the vocabulary was wrong, which took longer to see
than if the measurement had simply been broken.
Every bound in the tool was set by measuring eleven shots of real material and asserting that nothing fires. Not because real footage is well behaved, but because the first thing anyone does with a noisy alarm is stop reading it, so a threshold that flags genuine material has negative value.
| metric | real max | soft at | broken at | headroom |
|---|---|---|---|---|
incoherence | 0.00117 | 0.004 | 0.012 | 3.4x |
jerk | 1.78 | 1.9 | 2.5 | 1.07x |
closure | 0.0017 | 0.02 | 0.06 | 11.8x |
Across the whole real set exactly one finding fires, a soft on a stage camera that
genuinely reframed and came back, which is a true description of what that camera did.
soft never fails a gate; it exists so a real camera with a person walking through it
can be reported without being blocked.
Two things in that table are worth saying out loud rather than rounding off. jerk
has 7% of headroom, so a heavily operated long-lens shot could plausibly report
soft jerk one day. And the broken bound for jerk sits at 2.5, where
nothing in the set reaches, which means it has no positive control. It is a
conservative bound, not a validated one, and a bound with no control behind it is a guess with a
comment next to it. Writing that down is the only thing that stops it from being quoted later as
though it were measured.
All three were found by pointing the tool at real footage rather than at fixtures, which is the only part of this process I would call non-negotiable. Fixtures tell you the code does what you wrote. Real material tells you what you wrote was the wrong thing.
Two of the three shared a single root cause: a ratio whose denominator is legitimately
zero on a static shot. A locked-off tripod was reported as breathing 14.8x,
because net scale change on a static shot is genuinely zero and sub-pixel measurement noise divided
by nothing is a large number. Sub-pixel jitter was reported as wander for the same
reason, path length over a near-zero net. The fix in both cases is a floor on absolute travel, so
the ratio is only consulted once something has visibly moved.
Then I applied that fix inconsistently, which was its own small lesson.
wander kept an extra condition requiring the net displacement to be below a floor, and
that condition made it miss a control which travelled 0.74 of a frame width and ended 0.047 from
where it started, which is precisely the fault wander is named after. The ratio plus a travel gate
is sufficient. The leftover condition was redundant and harmful, and it survived review because it
looked like part of the fix.
The third was a 20-frame fragment between two cuts reported broken on
jerk, a third derivative estimated from 19 samples. Shape is no longer judged below 16
frame pairs.
None of those three is the one I think about. closure is the tool's independent
check: the per-frame chain says where frame k ended up by summing every step, and measuring
frame 0 against frame k directly asks the same question using none of those steps. Where
both are trustworthy they have to agree, which makes it a genuinely separate route to the same number
rather than a second opinion from the same code.
The first version believed any anchor it could compute. Once frame 0 and frame k barely overlap, the direct correlation does not fail. It returns a confident wrong answer. Measured on one clip at five surviving tiles and a fit residual a thousand times the clip's own, it reported +0.003 where the truth was +0.270. Taken at face value that is a catastrophic disagreement between the two routes, reported on a clip that is in fact perfect.
A false positive wastes somebody's afternoon. This one would have told a person their good footage was broken, using the tool's most authoritative-sounding check, and the evidence that the measurement was worthless was sitting right there in the tile count. An anchor now has to earn the right to contradict the chain, and one that cannot is skipped and reported as not measurable. Cannot measure must never be reported as measured bad, and the failure mode is not that your tool is wrong, it is that your tool is wrong in a confident voice.
Here is the thing I would want pointed out if I were reading this. shotdrift exists because of generative video, and there is no generated video in its calibration set. Every clip is either real camera footage or a synthetic control with a known fault. So the bounds are validated for two claims, that they do not flag reality and that they name faults I deliberately built, and they are unvalidated against the specific failure modes of any particular generator. Eleven real shots is also a small sample, all from live-event cameras: no drone, no gimbal, no handheld documentary, no film scan.
That is a real gap and it does not undermine the part that works, because the part that works is not a verdict about authenticity. It is a measurement of a camera path, plus a comparison against a move you declared. Both halves hold regardless of what produced the pixels, which is also why the same command runs on footage from a camera, a render, or somebody else's demo reel.
The general shape, which is the only thing here I would carry to the next tool: the check I was most confident about before measuring is the one that died, and it died because reality is noisier than the defect. That is not a special property of video. Any time you plan to detect badness by measuring deviation, the honest first experiment is to measure how much the good material deviates, and be prepared for the answer to be "more". When it is, the fix is usually not a better threshold. It is a better question, and the better question nearly always involves somebody's declared intent, because intent is the only thing that tells you which deviations were supposed to be there.