The most honest thing about the Jebadiah v2.1 release is the regression table. The creator of these open models, which treat decision-making as a closed-set scoring problem instead of text generation, published the exact points where performance dropped, Tools fell 2.10 on the 27B and 0.84 on the 9B, with the When2Call sub-score sliding more than six points. That kind of transparency is rare, and it tells us something important about where the field is headed. We already know from Android Bench 2.0 Measures AI Agents on Complex, Multi-Step Tasks that evaluating agents on multi-step actions is a genuinely hard problem. And when AI models stumble when spreadsheet rules demand creative combination, the lesson is the same: benchmarks expose weaknesses, and the best response is to publish them, not hide them.
The core idea here is worth paying attention to. Most large language models generate tokens until they produce an answer, then someone has to parse that output and hope the sampling didn't introduce noise. Jebadiah skips all of that. It takes structured context and a multiple-choice question, then outputs probabilities directly from the logits. No generation step, no sampling variance, no parsing errors. For a narrow but common class of decisions, classification, routing, tool selection, this approach is faster and more deterministic than anything built on autoregressive text. The v2.1 results show gains across Knowledge, Language, Retrieval, and Arts on both the 27B and 9B variants. Those are real improvements in areas that matter for everyday data work.
But the tool-use regression is the detail to watch. When2Call dropped sharply, and the creator admits they haven't found the cause yet. That is not a failure; it is an invitation. If you are building with these models, you now know exactly where to stress-test your own workflows. The quantization records are also unusually thorough: 259 out of 260 questions matched on the 27B Q8_0 build, and the per-question records are published. That level of precision lets you decide for yourself whether the trade-off is acceptable for your use case. Contamination scanning was done with a text-match scanner, which the creator notes would miss paraphrased overlap. That is an honest limitation, not a hidden one.
The practical takeaway is straightforward: for structured decision tasks where speed and reproducibility matter more than creative generation, this architecture deserves a real look. The weights are Apache-2.0, the code is public, and the evaluation is open to critique. The regression on tools is a loose end, and the creator has asked for help finding it. That is the kind of open problem the community should actually engage with, rather than waiting for a polished commercial release. Watch whether the When2Call issue gets resolved, because if it does, the case for skipping text generation in decision models becomes much harder to ignore.