Independent research has always carried a peculiar kind of weight: no institutional safety net, no lab brand to soften the landing. So when a solo researcher in rural Manitoba submits a 55-model, 22,000-judgment evaluation to arXiv, the first instinct is to check the math. The numbers hold. The Multivac extends the Panel of LLM Judges into a fully symmetric matrix, and the findings are not just statistically interesting; they are the kind of result that should make every frontier lab pause before quoting a leaderboard slot as if it were a trophy. Six different models top nine category pools. The most frequent first-place finisher ranks sixteenth by mean score. That is not a quibble. That is a structural critique of how we currently rank intelligence.
What makes this work useful is not the novelty of the idea, because symmetric peer evaluation has been discussed. It is the scale and the refusal to let the method hide behind aggregate gloss. The same-family bias appears in all eight families tested, with a previously unreported negative bias in Mistral and Google models. That is a concrete, falsifiable observation about how models judge their own relatives, and it has direct implications for any evaluation that mixes model families without controlling for it. The top four frontier models are pairwise indistinguishable by bootstrap confidence intervals. In plain terms, the difference between first and fourth place on a typical public leaderboard is noise, and the paper says so with data attached.
For the working practitioner, the practical takeaway is not that leaderboards are worthless. It is that they are blunt instruments, and this work gives you a sharper one. The code evaluation disagreement being roughly double that of meta-alignment tells you where the field's own uncertainty lives, and it is not where the marketing materials suggest. The full release under MIT license, with 27,540 judgments and all prompts, means this is not a locked-room study. Anyone can rerun it, test the categories, or extend the matrix. That is the standard the field should expect, and it is coming from a single author in rural Manitoba rather than a well-funded lab.
The endorsement request is modest, and that is part of why it deserves attention. Thirty seconds of an eligible author's time to validate a first submission that has already done the hard work. No paywall, no press release, just a request for a look. If the methodology holds up under review, this becomes a reference point for how independent evaluation can be done with rigor. If it does not, the open release means the flaws will be found quickly. Either way, the field moves forward. The concrete ask is simple: if you have the eligibility, use the code S33JQD. The data is already public. The contribution is already made. The only missing piece is a signature from someone willing to say that independent research deserves a seat at the table.