The gap this user uncovered tells us something uncomfortable about how we measure progress in AI. When a vision-language model achieves perfect accuracy on a multi-step reasoning task simply because you give it four options, but cannot produce the same answer on its own, the benchmark is not testing reasoning. It is testing recognition. And that distinction matters for anyone who relies on these tools for real work.
The user designed questions with no options, just a ground truth answer, and the models failed. Then they presented the same questions with four choices, and the models hit 100 percent. This is not a quirk. It reveals a fundamental limitation in how these models process long video content. They can match patterns. They can pick the correct answer from a shortlist. But they cannot yet construct a chain of logic across multiple steps without the scaffolding of predetermined options. For anyone building workflows around video analysis, reviewing security footage, auditing training videos, extracting insights from recorded meetings, this means the current generation of VLMs is not ready for the open-ended tasks you probably need them to handle.
The datasets we celebrate, like Video-MME and MLVU, focus on ordering, counting, and reasoning within defined categories. They are useful for comparing models, but they do not simulate the messy, multi-step questions that arise in practice. The user's experiment is a stress test that the benchmarks never ran. It shows that when you remove the training wheels of multiple-choice format, the models lose their footing. This should change how we evaluate progress. A model that scores high on existing benchmarks but cannot answer a simple open-ended question about the same video is not truly understanding what it sees.
Our position is straightforward: benchmarks should measure what users actually need, not what models can easily game. The community should design evaluation sets that require open-ended, multi-step reasoning from the start. Until then, treat 100 percent accuracy on multiple-choice video benchmarks with skepticism. Ask yourself whether the model could produce that answer without the list. For now, the evidence suggests it probably cannot.