The hand tracker saw the human hand approach the socket. Then, at the exact moment the plug needed to seat itself, the tracker blinked. It returned a moment later, the connector already in place. The video looks fine. The aggregate recall score is high. But the data now has a hole precisely where the physics of contact turned into success or failure. That gap is not a minor artifact; it is the entire point of the demonstration. For anyone building imitation learning pipelines from human motion data, this is the problem that aggregate metrics hide.
We have written about how Match strain lookup errors with an AI-powered spreadsheet approach can turn a messy engineering problem into a structured one, but structure only helps if you know where the mess is. The same principle applies here. A pose error calculated only on frames where tracking succeeded will report a fine result while silently discarding the failure. The MEgoVista evaluation protocol that assigns an error to missed detections instead of excluding them is the right instinct, but it still collapses everything into a single number per episode. That number tells you that something is missing, but not whether it is the approach, the contact, or the withdrawal. For manipulation data, the difference between those phases is the difference between a usable demonstration and a deceptive one.
This is a concrete design constraint, not an abstract debate. If you are building a dataset for robot learning, you need to decide whether an episode with a missing contact-phase label is still usable. The answer depends on what you are teaching. For a task like Nested lookup logic made simple for multi-row pivot table data, a missing intermediate step might still yield the right final value. For a robot learning to insert a plug, the contact phase is the mechanics of the task. An episode that skips it is not a partial demonstration; it is a demonstration of approach without resolution. The robot learns to reach the socket, but not to seat the connection. That is a subtle failure that aggregate recall will never surface.
The specific takeaway is this: when evaluating hand tracking for manipulation, do not accept episode-level F1 scores as sufficient. Demand coverage broken down by approach, contact, and withdrawal, and ask for the length of the longest consecutive gap during contact. If the tracker cannot deliver that breakdown, the data is not ready for training. A robot that learns from a gap learns to ignore the hardest part of the task. That is not a training strategy; it is a blind spot.
