The long-term memory benchmark landscape has a measurement problem, and it is worse than most teams realize. After auditing LoCoMo, we found that 6.4% of its answer key is wrong, hallucinated car models that never appeared in the source text, incorrect date arithmetic that penalizes correct reasoning, and speaker attribution errors that reward inaccurate tracking. A perfect memory system cannot score higher than 93.6% on this benchmark, and the LLM judge accepts 63% of intentionally wrong answers. When a benchmark cannot distinguish a correct answer from a plausible wrong one, the scores published in tables are not measuring memory. They are measuring noise.
This matters because teams are still submitting new scores on LoCoMo as of March 2026, treating it as a gold standard. The reality is that the benchmark rewards the failure mode of weak retrieval, locating the right conversation but extracting nothing specific, because vague answers pass the judge nearly two-thirds of the time. If your system finds the correct chat but returns "a red sports car" instead of the specific model, LoCoMo calls that a win. That is not a memory test. LongMemEval-S has a different but equally fundamental flaw: its entire corpus fits in a 115K-token context window, which most current models already handle. Mastra's research showed a full-context baseline scoring 60% with gpt-4o, meaning the benchmark measures context window management rather than long-term retrieval. As models grow to support millions of tokens, that baseline will climb and the benchmark will lose its ability to discriminate entirely.
The community needs to face a hard requirement: corpus size must exceed context windows. BEAM's approach with conversations up to 10 million tokens moves in the right direction, but it introduces its own challenges. More importantly, evaluation pipelines must be standardized or fully disclosed, ingestion methods, embedding models, judge prompts, and standard deviations. Without that, cross-system comparisons are not meaningful. Ground truth must be verified adversarially, and judge models must reflect current capabilities. Using gpt-4o-mini as a judge introduces a ceiling on scoring precision that makes small score differences uninterpretable. LoCoMo-Plus inherits all 99 errors from the original benchmark while adding a genuinely interesting cognitive category, but it does not fix the evaluation infrastructure.
For teams building memory systems, the practical takeaway is this: do not benchmark against noise. Run your own adversarial validation on any benchmark you adopt. Verify the answer key against source conversations. Test your judge with intentionally wrong answers to understand its failure modes. And if the full test corpus fits in a single context window, you are measuring retrieval within a conversation, not memory across time. The long-term memory evaluation problem is genuinely hard, it sits at the intersection of retrieval, reasoning, temporal understanding, and knowledge integration, but the first step is admitting that most current benchmarks do not measure what they claim.