The coalescing nobody flagged
Exit code 4 came back on the first run where the meter was actually allowed to say no. Before that, every run of the tracking harness had exited 0, and every 0 had a number attached: translation error in pixels, lag in milliseconds, comfortably inside the budget I’d written down before I ran anything. The 0 was wrong. It took an adversarial read of the meter itself to prove it.
The harness is three Win32 processes on one virtual desktop. Target is a plain window that gets moved around, by script or by hand. Tracker is the layered overlay under test, a strip that’s supposed to sit flush above the target no matter where it goes. Meter is the judge, and it does not call GetWindowRect on anything. It uses DXGI Desktop Duplication to capture the composited desktop, hunts for four 16x16 fiducial markers, two on the target, two on the overlay, and computes translation error, edge error and size error from where the markers actually landed on the glass. The whole point was to stop trusting Win32 coordinates and start trusting pixels.
Where the first matrix went wrong
The meter’s acquire loop originally did three things on one thread: call AcquireNextFrame, map the staging texture, and once a second, sample the tracker’s and target’s process stats for the CPU/GC/memory budget. The first matrix run came back clean, and also came back with a pattern I almost let slide: every frame the log marked coalesced landed on an integer-second boundary. Not near it, on it, every time.
A blocked capture thread produces exactly that pattern. The once-a-second stat sample was long enough to make AcquireNextFrame miss the next composited frame, and DXGI’s response to a slow consumer is silence: no error, no missed-frame count, nothing in the return value that flags it, it just hands you the next available frame with whatever happened in between folded in. I split the meter into three threads instead of one: acquire-only, a worker that maps (D3D11_MAP_FLAG_DO_NOT_WAIT, retried) and detects, and a separate 1 Hz sampler, so nothing but frame acquisition ever touches the acquire thread. That killed the boundary artifact. I treated it as the fix.
It fixed the case where I already knew to look for it.
The bias runs the wrong way
Desktop Duplication doesn’t queue frames. AcquireNextFrame gives you the current composited state, and if you were too slow to ask for the previous one, it quietly folds the missed updates into whatever it hands you next. There’s no counter for that. The only trace is that the frame you got spans a longer interval than the nominal refresh, and unless you’re checking against something outside the meter’s own clock, you don’t know it happened.
A review pass on the meter’s methodology is what put the real failure mode in front of me: this isn’t random jitter, it’s a bias that points in exactly the wrong direction. Coverage is worst precisely when the target is moving fastest, because a moving target generates composited updates at a higher rate, which leaves the worker thread less time per frame to map and detect before the next one lands. Detection under motion is also more expensive, since more of the region is plausibly a marker mid-transit and the tolerance search does more comparisons. So the frames most likely to get silently coalesced away are the ones from the segments where the tracker is most likely to be lagging, and the frames that survive to get measured skew toward the calmer stretches where the overlay already had time to settle. Average the pixel error over whatever frames.csv happens to contain and you get a number that looks great, right up until you watch the overlay on screen and see it stutter through a resize.
I had built an instrument that would report a near-perfect score for exactly the failure it existed to catch.
Making coverage a pass condition
The fix wasn’t a faster meter. Detection cost scales with marker count and search tolerance, not with how badly I wanted the acquire thread to keep up, and I wasn’t going to win a race against the compositor. The fix was making the instrument refuse to answer when it couldn’t see.
Meter now takes --target-log and --tracker-log, the same per-move logs Target and Tracker already write for their own diagnostics: kind 1 for a commanded move, kind 2 for WM_WINDOWPOSCHANGED, and so on down the list. Every gap in the captured-frame sequence gets checked against both. The target log gives the commanded velocity and amplitude during the gap, so the meter can compute the worst-case position error that gap could be hiding, and the rule is that every commanded move has to show up in a captured frame within 2.5 composited intervals or the gap gets charged against that bound. Without the tracker log, the meter doesn’t get to assume the overlay held still either; it assumes the overlay was frozen for the whole gap, the worst case, not the convenient one.
That bound is coverage’s teeth. It isn’t a capture percentage you eyeball and call close enough. For every stretch the meter didn’t see, it asks what’s the largest error that gap could be hiding, and whether that worst case still clears tolerance. If it doesn’t, the run doesn’t get to report a pixel number at all. Meter exits 0 when the lag and jitter claim is actually supported by coverage, 4 when it isn’t, 2 on setup failure. Exit 4 means the meter declined to answer.
The first run under the new rule came back exit 4. The tracker hadn’t gotten worse, it hadn’t been touched, but the honest accounting of what the meter couldn’t see, bounded against what the target log said was happening during those gaps, exceeded the budget I’d written down before any of this ran. The 0 I’d been looking at for two matrix runs had never been a real 0. It was an average over a sample that had already had its worst evidence quietly removed before I ever saw it.