Docs · Method

Benchmark protocols

What existing benchmarks measure, what our runners actually exercise, and why the two are reported separately.

Coverage is not compatibility

A benchmark measures certain behaviours. Our runner exercises certain interfaces. Those are two different limits, and neither means a system is incompatible.

Existing benchmarks can test Mnemosyne. They just do not cover all of what it does.

What the upstream protocols cover

LongMemEvalExtraction, multi-session reasoning, updates, temporal reasoning, abstention. Explicitly supports custom systems.
LoCoMoConversational QA, event summarisation, multimodal dialogue. A QA-only run covers part of the release.
MemoryAgentBenchRetrieval, test-time learning, long-range understanding, conflict resolution.
HippoRAG suiteMulti-hop retrieval over MuSiQue, 2Wiki and HotpotQA.

Our run uses the original LongMemEval, not LongMemEval-V2.

What our runner actually does

It loads each question’s history, searches it, and scores which sessions came back. That is Recall@5 and nDCG@5 only, not answer quality.

It does not test forgetting, deletion, scheduled actions or isolation. Those need their own workloads.

Why LoCoMo is never headlined

LoCoMo is the most quoted memory benchmark and the most disputed: its answer key has known errors and its judge is lenient, so scores swing by several points depending on who runs it.

Three tracks, separate conclusions

Official upstreamThe original protocol, inputs and scoring, unchanged.
Enhanced successorSeparately versioned tests with disclosed differences and controls.
DevelopmentLocal fixtures and conformance tests, never substituted for official results.
Internal onlyLongMemEval-QA under the current publication policy.

A low score is still a real result. Untested features do not cancel it.