FLEET

Reward-aware search replaces blind sampling in best-of-N generation

Repetitive sampling in best-of-N generation is basically a blind search, even though most tasks using it are built around reward maximization.

3 min readMachine Learning

For too long, the dominant approach to improving AI model outputs has been brute force: generate many samples, pick the best one, and hope the process gets smarter over time. It works, but it is wasteful. The FLEET algorithm, introduced by a team including one of our readers, finally replaces that blind sampling with something far more intelligent. Instead of repeatedly guessing in the dark, FLEET learns from each attempt by attributing external rewards to specific tokens and using that information to guide the next generation. This is not a minor efficiency tweak; it is a fundamental shift in how we think about iterative generation.

The core insight is elegantly simple. When a model generates a completion, it produces a sequence of tokens, each with varying levels of uncertainty. FLEET identifies the tokens where the model is most unsure, those with high entropy and varentropy, and treats them as branching points. It stores these states in a vector store, mapping them to metadata about past rewards and transitions. On subsequent runs, instead of starting from scratch, the algorithm retrieves this history and uses a modified Monte Carlo Tree Search to penalize paths that have proven suboptimal. The results speak for themselves: on GSM8K, FLEET matched the performance of standard best-of-N sampling using half the iterations. On LiveCodeBench v6, it boosted the score from 0.59 to 0.69 under the same budget, reaching the baseline with just nine iterations instead of thirty-two.

This matters because it challenges a long-held assumption in the field. As we have seen in New study finds AI models bend facts for verified sources, models often stubbornly cling to errors even when presented with correct information. That behavior is a symptom of a system that lacks awareness of its own past mistakes. FLEET takes a different path: it treats each generation as a learning opportunity, not a throwaway experiment. The metadata store can even be preserved as a prior for other tasks or used to enrich supervised fine-tuning, making the investment in search pay dividends across multiple use cases. This is the kind of practical, feedback-driven design that moves AI from a black box to a transparent tool.

The most concrete takeaway here is that the era of blind sampling is ending. For anyone building applications that rely on repeated generation, whether for code synthesis, mathematical reasoning, or creative writing, FLEET offers a blueprint for doing more with less. The sequential execution is not even required; the lookup table can be passed in as a prior, meaning the efficiency gains are accessible without architectural overhauls. The open question is how well this approach scales to larger models and more diverse tasks, but the principle is sound. We should expect reward-aware search to become the new baseline, not the exception. And that is a future worth exploring.

From Machine Learning

I am one of the authors of FLEET - an algorithm that enhances Best-of-N generation by attributing external rewards to particular tokens and then uses MCTS to adjust logits during the next run.

Read the original at Machine Learning