Docs · Method
Benchmark protocols
What existing benchmarks measure, what our runners actually exercise, and why the two are reported separately.
Coverage is not compatibility
A benchmark measures certain behaviours. Our runner exercises certain interfaces. Those are two different limits, and neither means a system is incompatible.
What the upstream protocols cover
Our run uses the original LongMemEval, not LongMemEval-V2.
What our runner actually does
It loads each question’s history, searches it, and scores which sessions came back. That is Recall@5 and nDCG@5 only, not answer quality.
It does not test forgetting, deletion, scheduled actions or isolation. Those need their own workloads.
Why LoCoMo is never headlined
LoCoMo is the most quoted memory benchmark and the most disputed: its answer key has known errors and its judge is lenient, so scores swing by several points depending on who runs it.
Three tracks, separate conclusions
A low score is still a real result. Untested features do not cancel it.