There is a quiet power in showing your work, especially the unfinished parts. The open-source Clash Royale simulator and its live training demo, shared by a developer on Reddit, do exactly that. They put a reinforcement learning agent on display as it learns a single defensive decision: where to place one card and how long to wait. The reward is simple, tower damage prevented, and the optimum is known, found by brute force. So the gap between what the policy achieves and what is possible is right there on the chart, visible to anyone who clicks the link. This is not a polished product. It is a miniature of a much larger problem, and the creator is upfront about that. The honesty is the point.
For our readers who have been following the Build AI from the ground up with 523 hands-on lessons, now in portable books curriculum or the emergent tactics seen in Teaching AI to Brawl: Emergent Tactics From Reward-Shaped Sparring, this demo offers something different. It is not about scale or surprise. It is about making the mechanics of training visible. The policy has only 5,629 parameters. It trains with REINFORCE in plain JavaScript, using hand-written gradients. Every rollout runs in a C++ engine compiled to WebAssembly, and the deploy pipeline checks that the WASM build matches the native engine exactly. That level of engineering discipline, applied to a toy problem, is what makes this worth studying. It shows that understanding the loop matters more than the size of the model.
The most instructive detail is the local optimum problem. Against a Giant vs Cannon matchup, a lane placement worth about 75% of the best answer acts as a trap. With a constant entropy coefficient, five of six runs stayed there. A linear anneal from 0.1 to 0.005 over 10,000 tries reduced that to one of six. That is a concrete, quotable takeaway: entropy annealing is not a theoretical nicety; it is the difference between getting stuck and finding the better answer. The demo also withholds one pairing, Battle Ram vs Valkyrie, because no setting got past 55% of the optimum. That is not a failure to hide. It is an honest boundary, and it invites the question of what would be needed to cross it.
What would we tell a reader who asked about this? Look at the chart. Watch the policy improve. Notice where it stalls. Then ask yourself whether your own data workflows have visible feedback loops, or whether you are operating on faith. This demo is a reminder that even in a toy problem, the path from random to competent is not smooth, and the best answer is not always the one the algorithm finds first. The open question is whether the same principles, small policies, hand-written gradients, honest baselines, can scale to the full 4-card hand with elixir and recurrent PPO. The creator does not claim they can. But the discipline required to build this miniature is exactly what makes scaling possible.