LLM

LLM-guided evolution unlocks new best solutions for circle packing

Ten of the best-known circle-packing solutions on the Packomania csqv benchmark just got better, and an LLM did the improving.

4 min readMachine Learning

The most striking number in the Packomania result isn't the 5.4% improvement or even the 10 new best-known solutions. It's the $27.72. For less than the cost of a decent dinner, someone used an LLM to evolve an optimization algorithm that outperformed decades of human expertise on a notoriously hard benchmark. That's not just a clever hack. It's a quiet signal that the way we build algorithms is about to change, and the change won't come from bigger models or more compute. It will come from rethinking the loop between human intent and machine iteration.

The approach here is deceptively simple. Instead of asking the LLM to solve the circle-packing problem directly, the author let it propose changes to a basic seed solver, scored each candidate against an independent verifier, kept what worked, and discarded the rest. This is evolutionary computation with a language model as the mutation operator. What makes it compelling is the division of labor. The LLM brings the creativity, the verifier brings the ground truth, and the human brings the judgment about when to stop. That last part matters, because the author explicitly flags the plateau-detection stopping rule as the piece they'd most want critiqued. They're right to flag it. The algorithm improved results in 15 iterations, but without a principled stopping rule, you're left wondering whether a few more rounds would have yielded even more gains, or whether the system was just starting to overfit to the benchmark's quirks.

This is where the story connects to the broader challenges of deploying AI in the real world. In Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges, we saw how models that work in the lab often stumble when they hit the messy constraints of production. The same lesson applies here, but in reverse. The LLM didn't need to understand the geometry of circles to improve the packing. It just needed a clear feedback loop and a way to express small, testable changes. That's a powerful reminder that the bottleneck in AI-assisted discovery isn't always the model's intelligence. It's often the design of the evaluation process around it.

The other connection is to how we think about tools themselves. In Explore the Forrester Function: Beyond Mathematics, a Tool for Machine Learning, we explored how abstract functions can serve as practical testbeds for optimization methods. The Packomania result takes that idea further. Here, the function isn't just a testbed. It's a real, unsolved benchmark with decades of history. The fact that an LLM-driven loop could beat the best-known solutions for 10 values of N, with a total cost of under thirty dollars, suggests that the barrier to entry for algorithmic discovery is collapsing. You no longer need a specialized research team to push the state of the art. You need a clear problem statement, a verifier, and the willingness to let a language model iterate.

What would we tell a reader who asks whether they should try this at home? Start with a problem that has a cheap, reliable scoring function. Don't aim for a breakthrough. Aim for a small, honest improvement. The real lesson from this paper isn't that LLMs are magic. It's that the loop matters more than the model. The human's job is to define what "better" means and to know when to stop. The LLM's job is to explore the space of possible changes without getting attached to any one idea. That's a collaboration worth paying attention to, especially as the cost of running these loops continues to fall. The next time you see a stubborn optimization problem, ask yourself whether you need a better algorithm or just a better way to evolve the one you have. That question is worth more than any single solution.

From Machine Learning

I used an LLM to iteratively evolve an optimization algorithm rather than solve the packing directly. Starting from a simple seed solver, the LLM proposes algorithmic changes guided by a scoreboard of results and a history of prior attempts, and each candidate is scored by an independent verifier so improvements are kept and failures discarded. On the Packomania csqv benchmark it improved the best-known sum-of-radii for 10 values of N from 101 to 114, by 2.4 to 5.4%, in 15 iterations. Total LLM cost was $27.72. Packomania accepted the results independently.

Code + solutions: github.com/ucsandman/discovery-loop

Read the original at Machine Learning