Real Robot AI Faces Its First Honest Benchmark: And the Results Are Humble

Introducing PhAIL (phail.ai), an open benchmark designed to assess robot AI performance on real hardware, focusing on practical metrics rather than simulations. After a year of investigation into VLA models’…

3 min readMachine Learning

The numbers are in, and they are not flattering. After a year of independent testing on real warehouse hardware, the gap between what these models promise and what they deliver is not a gap at all; it is a chasm. The benchmark, built by a developer who was tired of marketing masquerading as metrics, shows the best vision-language-action model managing 65 units per hour against a human-operated teleop baseline of 330. That is not a slow start. That is a fundamental difference in capability, and it is the most honest data we have seen in this space.

For anyone evaluating automation, this is the reality check you have been waiting for. The mean time between failures, four minutes for the top performer, tells the real story. A robot that needs a full-time babysitter is not autonomous, it is an expensive hobby. The economic case for deployment collapses when you factor in the human supervisor required to catch every mistake. This is not about the technology being new; it is about the technology being unreliable in the exact conditions where it is supposed to create value. The developer's decision to measure UPH and MTBF, the metrics operations managers actually use, exposes how much of the prior hype was built on carefully selected demo clips.

What makes this work stand out is the refusal to hide behind favorable conditions. The evaluation was blind, the data is public, and the toolkit is open for anyone to verify or challenge. This is how you build trust in a field drowning in exaggerated claims. The developer is not asking you to take their word for it; they are inviting you to run the tests yourself. That is the right approach, and it sets a standard that other research groups should be forced to meet.

The path forward is not to abandon these models but to push them toward the reliability threshold that makes them viable. The teleop baseline of 330 UPH is the target, and it is a high one. Until a model can operate for hours without intervention, the term autonomous will remain aspirational. The next step is clear: submit your own checkpoint, challenge the leaderboard, and prove that the technology can close the gap. That is the only way this field earns its keep, and this benchmark gives you the tools to try.

From Machine Learning

I spent the last year trying to answer a simple question: how good are VLA models on real commercial tasks? Not demos, not simulation, not success rates on 10 tries. Actual production metrics on real hardware.

I couldn't find honest numbers anywhere, so I built a benchmark.

Read the original at Machine Learning