← Syamjith NK

6 October 2026 · Syamjith NK

I set out to detect melting AI video. Real footage measured worse.

You hand a generative video model a camera move and it is under no obligation to deliver one. The usual check is a person watching sixty clips and forming an impression. So I wrote shotdrift, which measures the camera path out of the pixels instead: a grid of tiles tracked by phase correlation, one similarity transform fitted to all of them, and pan, zoom and roll falling out of the fit frame by frame. No metadata, no depth, no solve, because none of those come with the mp4.

The headline feature was going to be the melting-background tell. Everyone who has looked at generated video knows it: geometry that invents itself between one frame and the next, a wall that reorganises while the camera passes it. A tool that fits one rigid camera to a scene has, in principle, exactly the right residual sitting there to catch it, because whatever the camera model cannot explain is the other half of the product.

That feature is not in the tool, and the reason is a number I did not expect.

Real footage changes structure three times more than a literal dissolve

The measurement has to be comparable across clips, so the comparison frame is chosen adaptively: far enough away that the amount of warp between the two frames is the same every time, which keeps the resampling residue from being a function of how fast the shot happens to move. Then you measure what is left over.

The validation material is real single-camera footage from live events, 720p and 1080p, 30p and 50p, locked-off and operated, against synthetic controls built by cropping a textured plate so the fault in each one is known. The worst control I could build for this check was a cross-dissolve between two entirely different worlds, which is the limit case of invented geometry: every pixel in the frame is simultaneously two unrelated places.

clipspan residual
real: stage camera, a 20-frame segment1.062
real: stage camera, a 25-frame segment0.867
real: stage camera, operated slow zoom0.781
real: panel with a vision mix0.423
control: cross-dissolve between two different worlds0.358
control: clean push-in0.114
control: clean pan0.107

Read the order. Real footage is at the top. The thing the check exists to catch sits below all of it, closer to a clean synthetic push-in than to any real camera. The ordering is not marginal and it is not noise: it is reversed, by a factor of three.

The explanation is obvious once you see the number, and it is the reason no amount of tuning rescues this. Real scenes are full of things that are not the camera. People walk through them. LED walls behind a stage change their entire content on a cue. A vision mix cuts the picture out from under you. All of that is structural change, and there is more of it in a normal event clip than there is in geometry that melts. The melting tell is a thing your eye catches instantly and a thing this measurement cannot distinguish from Tuesday.

Removed, not tuned, and the blind spot is pinned as a test

There is no threshold in either direction, so there is no version of this check worth shipping. It is deleted. What is in the repository instead is a calibration document that records the distribution, states the failure, and keeps the dissolve control as a test.

That last part matters more than the deletion. The dissolve under a locked-off camera comes back clean, and a test asserts that it keeps coming back clean. It looks backwards to pin your own blind spot as a passing test, but the alternative is worse: a known gap that nobody wrote down slowly becomes a false positive the first time somebody tightens an unrelated bound, and then the tool is quietly flagging dissolves for a reason nobody can reconstruct. A stated limitation with a test behind it stays a limitation. An unstated one turns into a mystery.

The residual did not go to waste, either. Per-frame structural change is a poor melt detector and an excellent cut detector, which is the job it now does: real footage within a single shot reaches 0.36, and a transition that preserves some structure, a whip or a dissolve, spikes to between 1.7 and 8.6. Across a clean hard cut between unrelated shots nothing correlates at all, no camera can be fitted, and the value is NaN. The first version only tested for the spike, so it caught smeared transitions and walked straight past an ordinary cut.

Counting direction changes failed the same way

The second idea was reversals. If a camera move keeps changing its mind, something is wrong with it. This is the kind of metric that feels unarguable until you point it at material.

sourcereversals
real: stage camera, operated slow zoom27
real: panel with a vision mix16
control: a push-in that deliberately reverses twice3

An operator riding a zoom rocker reverses constantly, because that is what holding a shot on a live subject looks like from the inside. The real material scores nine times worse than the deliberate fault. Normalising by clip length makes it worse rather than better, since the long clips are the operated ones. Two checks in, and both of my instincts about what "wrong" looks like had been contradicted by the same property of real footage.

So reversals are still reported, labelled (measured, not judged), and nothing is gated on them.

What survived is intent, and that is the actual lesson

Both dead checks were asking an unanswerable question: is this video wrong? The question that is answerable is narrower and more useful. Did the move you asked for happen?

$ shotdrift pan_demo.mp4 --expect push-in

  camera path
    dominant move      pan
    pan                0.7159 of frame width
    zoom               1.000x
    reversals          0  (measured, not judged)
    incoherence        0.0000
    confidence         0.98

  asked for: push-in
    NOT HELD - asked for push in / dolly in; zoom moved +0.0001, under the 0.02
               floor - that move did not happen

  no findings: the motion is consistent with one physical camera.

That clip is a perfectly good shot. Smooth, coherent, one physical camera, nothing wrong with it at all. It is also not the shot that was ordered: it pans, and the push-in never happened. Both things are true and they are reported separately, because they are different questions and they have different fixes.

There are fourteen declared moves, and each is checked three ways, because the three failures are not the same failure: did it happen at all, was it the right way round, and was it held, or did a quarter of the travel run backwards. That last one is what the reversal count was reaching for all along, and it works here only because the declared move supplies the direction that makes "backwards" mean something. Without knowing what was asked for, a reversal is suspicious. With it, a reversal is measurable.

The signs are camera-relative, which produced the silliest bug in the project. A camera panning right makes the picture move left, so pan-right expects a negative picture-x. I wrote those the intuitive way round, and a correctly measured, perfectly clean pan reported "x went the other way". The measurement was right and the vocabulary was wrong, which took longer to see than if the measurement had simply been broken.

Silence on real footage is the requirement, and it costs you

Every bound in the tool was set by measuring eleven shots of real material and asserting that nothing fires. Not because real footage is well behaved, but because the first thing anyone does with a noisy alarm is stop reading it, so a threshold that flags genuine material has negative value.

metricreal maxsoft atbroken atheadroom
incoherence0.001170.0040.0123.4x
jerk1.781.92.51.07x
closure0.00170.020.0611.8x

Across the whole real set exactly one finding fires, a soft on a stage camera that genuinely reframed and came back, which is a true description of what that camera did. soft never fails a gate; it exists so a real camera with a person walking through it can be reported without being blocked.

Two things in that table are worth saying out loud rather than rounding off. jerk has 7% of headroom, so a heavily operated long-lens shot could plausibly report soft jerk one day. And the broken bound for jerk sits at 2.5, where nothing in the set reaches, which means it has no positive control. It is a conservative bound, not a validated one, and a bound with no control behind it is a guess with a comment next to it. Writing that down is the only thing that stops it from being quoted later as though it were measured.

Three false positives, and two of them were one mistake

All three were found by pointing the tool at real footage rather than at fixtures, which is the only part of this process I would call non-negotiable. Fixtures tell you the code does what you wrote. Real material tells you what you wrote was the wrong thing.

Two of the three shared a single root cause: a ratio whose denominator is legitimately zero on a static shot. A locked-off tripod was reported as breathing 14.8x, because net scale change on a static shot is genuinely zero and sub-pixel measurement noise divided by nothing is a large number. Sub-pixel jitter was reported as wander for the same reason, path length over a near-zero net. The fix in both cases is a floor on absolute travel, so the ratio is only consulted once something has visibly moved.

Then I applied that fix inconsistently, which was its own small lesson. wander kept an extra condition requiring the net displacement to be below a floor, and that condition made it miss a control which travelled 0.74 of a frame width and ended 0.047 from where it started, which is precisely the fault wander is named after. The ratio plus a travel gate is sufficient. The leftover condition was redundant and harmful, and it survived review because it looked like part of the fix.

The third was a 20-frame fragment between two cuts reported broken on jerk, a third derivative estimated from 19 samples. Shape is no longer judged below 16 frame pairs.

The inverse error was worse

None of those three is the one I think about. closure is the tool's independent check: the per-frame chain says where frame k ended up by summing every step, and measuring frame 0 against frame k directly asks the same question using none of those steps. Where both are trustworthy they have to agree, which makes it a genuinely separate route to the same number rather than a second opinion from the same code.

The first version believed any anchor it could compute. Once frame 0 and frame k barely overlap, the direct correlation does not fail. It returns a confident wrong answer. Measured on one clip at five surviving tiles and a fit residual a thousand times the clip's own, it reported +0.003 where the truth was +0.270. Taken at face value that is a catastrophic disagreement between the two routes, reported on a clip that is in fact perfect.

A false positive wastes somebody's afternoon. This one would have told a person their good footage was broken, using the tool's most authoritative-sounding check, and the evidence that the measurement was worthless was sitting right there in the tile count. An anchor now has to earn the right to contradict the chain, and one that cannot is skipped and reported as not measurable. Cannot measure must never be reported as measured bad, and the failure mode is not that your tool is wrong, it is that your tool is wrong in a confident voice.

The limitation I cannot test my way out of

Here is the thing I would want pointed out if I were reading this. shotdrift exists because of generative video, and there is no generated video in its calibration set. Every clip is either real camera footage or a synthetic control with a known fault. So the bounds are validated for two claims, that they do not flag reality and that they name faults I deliberately built, and they are unvalidated against the specific failure modes of any particular generator. Eleven real shots is also a small sample, all from live-event cameras: no drone, no gimbal, no handheld documentary, no film scan.

That is a real gap and it does not undermine the part that works, because the part that works is not a verdict about authenticity. It is a measurement of a camera path, plus a comparison against a move you declared. Both halves hold regardless of what produced the pixels, which is also why the same command runs on footage from a camera, a render, or somebody else's demo reel.

The general shape, which is the only thing here I would carry to the next tool: the check I was most confident about before measuring is the one that died, and it died because reality is noisier than the defect. That is not a special property of video. Any time you plan to detect badness by measuring deviation, the honest first experiment is to measure how much the good material deviates, and be prepared for the answer to be "more". When it is, the fix is usually not a better threshold. It is a better question, and the better question nearly always involves somebody's declared intent, because intent is the only thing that tells you which deviations were supposed to be there.