MemPalace's perfect score claims don't survive contact with its own benchmarks.

The recent launch of MemPalace, an open-source memory project, has sparked significant attention with claims of achieving "100% on LoCoMo" and a "perfect score" on LongMemEval.

3 min readMachine Learning

The numbers were never the story. The story is that MemPalace's own benchmarks file, buried under a week of viral tweets and a seven-thousand-star GitHub surge, quietly admits the headline claims don't survive contact with the methodology. The LoCoMo "100%" isn't a result; it's a top_k bypass that feeds the entire conversation into a reranker, turning retrieval into reading comprehension. The LongMemEval "perfect score" isn't a score at all; it's recall_any@5 on a retrieval-only task, with no answer generation and no judge. And the compression claim? The project's own data shows a 12.4-point quality drop on the same metric after compression. Lossless doesn't do that.

For you, the practical takeaway isn't that MemPalace is uniquely bad. It isn't. The field has been here before, Zep vs. Mem0, Letta's reproducibility critique, the Penfield audit of LoCoMo's answer key. What's different this time is that the disclaimers were sitting in the repository, in plain text, and the launch communication chose to strip them. That's not a technical failure. That's a decision. And it's the decision that matters, because it tells you something about how this project treats its own evidence. When a project's own documentation says "this is teaching to the test," and the tweet says "perfect score," you're not looking at a bug. You're looking at a communication strategy.

Here's what that means for your workflow. If you're evaluating memory systems for a real product, you cannot trust headline numbers from any project that doesn't ship its evaluation harness alongside its code. MemPalace does ship its harness, that's the one genuinely good instinct here, but it also ships a README that says one thing and a BENCHMARKS.md that says another. The gap between those two documents is where you'll find the actual performance. The same is true for every other project in this space. Run the evaluation yourself, on your own data, with your own judge. If the project won't give you the exact prompts, the exact splits, and the exact post-processing, assume the numbers are doing more work than the code.

The uncomfortable truth is that MemPalace is not an outlier. It's a stress test of what the field rewards. It got 1.5 million views and 7,000 stars because the claims were loud and the caveats were quiet. The fix isn't shaming this project, though a little accountability wouldn't hurt. The fix is making benchmark submissions as boring as code review. Until evaluation pipelines are standardized, adversarially validated, and mandatory to publish alongside results, we're going to keep seeing 100% scores that turn out to be 60% retrieval with extra steps. Read the BENCHMARKS.md. Ignore the tweet. And if a project won't show you its failure modes in the same font size as its successes, assume they're hiding something. That's not cynicism. That's just the cost of doing business in a field where the data is honest and the headlines aren't.

From Machine Learning

A new open-source memory project called MemPalace launched yesterday claiming "100% on LoCoMo" and "the first perfect score ever recorded on LongMemEval. 500/500 questions, every category at 100%." The launch tweet went viral reaching over 1.5 million views while the repository picked up over 7,000 GitHub stars in less than 24 hours.

Read the original at Machine Learning