AI Agent Gaming Tournament

AI Agents Face Off in UCLA's Open Tournament Across Four Games

On October 16, UCLA's Trustworthy AI Lab will put AI agents through their paces in an open tournament spanning Pokémon Showdown, Werewolf, Red Alert, and Honor of Kings.

3 min readMachine Learning

The most honest way to test an AI isn't a benchmark, it's a bluff. That's why UCLA's Trustworthy AI Lab tournament, where agents face off in Pokémon Showdown, Werewolf, Red Alert, and Honor of Kings, is more than a spectacle. It's a practical stress test for the systems we're already starting to rely on, and it deserves your attention before submissions close on the 13th.

What makes this competition compelling is the variety of failure modes it exposes. A spreadsheet model can't bluff its way through a game of Werewolf, and a chatbot won't survive Red Alert's real-time pressure. These games demand negotiation, deception, and split-second judgment, the messy skills that traditional benchmarks ignore. This feels like a natural evolution from the sharper decision-making models we saw with Cloudflare Clef gives AI agents a sharper decision-making model, which focus on picking the right action from a set of choices. UCLA's tournament goes further by testing whether agents can generate those choices under social and strategic pressure. For anyone building AI tools, this is the difference between a model that answers correctly and one that acts wisely when the rules aren't spelled out.

The practical appeal here is the low barrier to entry. You can bring your own agent and connect it through MCP, or use Oracle's prebuilt agents that just need instructions. That accessibility matters. It means a developer with a solid prompt can compete against a lab's custom system, and we all get to see what separates them. The $5,000 prize pool and backing from Oracle, Replit, and Matcherino add legitimacy, but the real payoff is observational. This is open research where the spectators get to watch the data being born. It also signals a shift in how we evaluate AI, moving away from static QA toward dynamic, adversarial environments. That's a trend worth tracking, especially as we see similar energy in how From ad spend to asset creation, Melius raises $20M to reimagine marketing tools is rethinking creative workflows, where AI isn't just optimizing but generating under constraints.

The specific detail to watch is Werewolf. It's the purest test of social reasoning, and it's the one game where a truly trustworthy AI should fail. A system that's honest and transparent will get eaten by the wolves. That tension, between building agents that are reliable and building agents that can deceive, is the open question this tournament forces us to confront. If UCLA's lab can measure that gap, they'll have given us something more useful than another leaderboard: a map of where our tools actually break. The submissions deadline is the 13th, so the field isn't set yet. The real question isn't who wins, but which agent learns to lie convincingly without breaking its own moral code.

From Machine Learning

UCLA’s Trustworthy AI Lab is running a tournament on October 16 where AI agents compete in Pokémon Showdown, Werewolf, Red Alert, and Honor of Kings. It’s open to everyone, including remote participants, with a $5,000 total prize pool secured thus far and support from Oracle, Replit, Matcherino, among other sponsors.

Read the original at Machine Learning