Explore how AI benchmarks reveal gaps in multi-step video reasoning.

In exploring long video understanding datasets like Video-MME, MLVU, and others, I've observed a significant gap in multi-step reasoning tasks.

3 min readMachine Learning

The gap this user uncovered tells us something uncomfortable about how we measure progress in AI. When a vision-language model achieves perfect accuracy on a multi-step reasoning task simply because you give it four options, but cannot produce the same answer on its own, the benchmark is not testing reasoning. It is testing recognition. And that distinction matters for anyone who relies on these tools for real work.

The user designed questions with no options, just a ground truth answer, and the models failed. Then they presented the same questions with four choices, and the models hit 100 percent. This is not a quirk. It reveals a fundamental limitation in how these models process long video content. They can match patterns. They can pick the correct answer from a shortlist. But they cannot yet construct a chain of logic across multiple steps without the scaffolding of predetermined options. For anyone building workflows around video analysis, reviewing security footage, auditing training videos, extracting insights from recorded meetings, this means the current generation of VLMs is not ready for the open-ended tasks you probably need them to handle.

The datasets we celebrate, like Video-MME and MLVU, focus on ordering, counting, and reasoning within defined categories. They are useful for comparing models, but they do not simulate the messy, multi-step questions that arise in practice. The user's experiment is a stress test that the benchmarks never ran. It shows that when you remove the training wheels of multiple-choice format, the models lose their footing. This should change how we evaluate progress. A model that scores high on existing benchmarks but cannot answer a simple open-ended question about the same video is not truly understanding what it sees.

Our position is straightforward: benchmarks should measure what users actually need, not what models can easily game. The community should design evaluation sets that require open-ended, multi-step reasoning from the start. Until then, treat 100 percent accuracy on multiple-choice video benchmarks with skepticism. Ask yourself whether the model could produce that answer without the list. For now, the evidence suggests it probably cannot.

From Machine Learning

I have extensively searched on long video understanding datasets such as Video-MME, MLVU, VideoBench, LongVideoBench and etc. What I have seen there these datasets are focused on different categories such dramas, films, TV shows, documentaries where focus on tasks like ordering, counting, reasoning and etc.

I feel that multi-step reasoning is less explored and then what i have did i designed the questions with no options just ground truth and asked the VLM to give me the answer but VLMs unable to give the answer. But when i give the 4 options then VLM achieves 100% accuracy.

Read the original at Machine Learning