Existing benchmarks can evaluate Mnemosyne, but a score does not describe its whole memory system. Coverage means which behaviors a protocol measures. Adapter compatibility means which system interfaces a runner actually exercises. Limited coverage is not evidence of incompatibility.
For Mnemosyne, the accurate limitation is partial measurement of the implemented memory lifecycle, not that conventional benchmarks cannot measure it. A question-answering test can measure the benefit of an internal memory mechanism when that mechanism is enabled during the run. It cannot establish the behavior of operations the run never invokes. Any missing integration belongs to our adapter status, not to an assumption that the benchmark is inherently incompatible.
LongMemEval explicitly supports custom systems: process timestamped histories and submit answers. Its original tasks cover extraction, multi-session reasoning, updates, temporal reasoning and abstention. LoCoMo covers conversational QA, event summarization and multimodal dialog generation; a QA-only run covers only part of that release. These are useful measurements, not tests of every operational property. These statements concern the original protocols, not every newer version or benchmark in this catalog.
Version matters: the LongMemEval repository now points to LongMemEval-V2. The coverage discussion here concerns the original protocol used by our current run; it is not a coverage assessment of V2.
Our current runner has a narrower view
The local LongMemEval retrieval runner batch-captures each question’s history into an isolated store, searches it and scores retrieved session IDs. It measures fractional Recall@5 and conventional nDCG@5, not generated-answer quality. These historical metrics differ from upstream LongMemEval’s all-evidence recall and DCG discount; they must not be presented as unchanged upstream scores. A separate versioned adapter now exercises the upstream formulas on a registered synthetic development suite. A full held-out run under that profile has not been completed. It does not explicitly drive scheduled consolidation or rehearsal over time, branch/merge workflows, recurring actions, deletion verification or tenant-isolation attacks. Those capabilities need their own exercised interfaces and evidence.
For example, answering a question about dates tests temporal reasoning. It does not by itself test whether a memory survives months of interference, whether scheduled rehearsal preserves it, or whether a deletion prevents it from resurfacing. Those are different behaviors requiring explicit workloads and checks. A retrieval score alone cannot establish them.
Mnemosyne already has implementations and regression tests for several of these behaviors, including branch/merge, recurring actions and protected-memory rehearsal. Their complete whole-memory evaluations remain unfinished. Inspect implementation evidence and its limits · See module status.
Match the memory interface and the retrieval budget
A benchmark may divide a document into large chunks while a memory system limits the total context returned by search. A token budget smaller than those chunks can exclude evidence before the answer model sees it. Matching a number called “tokens” is not enough: systems can use different tokenizers or estimates, and memory wrappers also consume space.
Our local MemoryAgentBench integration exposed this exact kind of mismatch with Mnemosyne’s default context budget. That run is a compatibility diagnostic, not a fair basis for ranking answer quality against the official multi-chunk retrieval setup. An expanded-budget experiment must retain the original result, declare the changed settings, and report resource use alongside quality.
The same rule applies to every participating system: verify what was actually stored, what search could return, and what the answer model received. Do not silently shrink the task, raise one system’s budget, drop empty retrievals, or attribute an adapter limitation to the upstream benchmark.
Fair evidence, including weaknesses
A low score remains a real result for the tested configuration; untested features do not cancel it or prove superiority. We must preserve upstream protocols for comparable results and disclose adapter limitations. Additional whole-memory tests belong in a separately identified track, with the same rules available to every participating system.
Three tracks, separate conclusions
Official upstream: the original protocol, inputs and scoring, unchanged. Enhanced successor: separately versioned tests with disclosed differences and controls. Development: local fixtures and conformance tests, never substituted for official results.
LongMemEval-QA is internal-only under the current publication policy. Retrieval and answer quality remain separate. No combined overall rank is implied.
Read the full slate and source audit