Official benchmark setup required. Current development checks and adapted results are not official benchmark runs. Official comparisons must use the upstream code, datasets, scoring and prescribed setup, with versions and deviations disclosed. Protocol status →
Memory benchmarks, with evidence.
Evidence for AI memory. Compare measured results and inspect the traces behind every number.
Results
| System | Benchmark | Metric | Result | Evidence |
|---|---|---|---|---|
| mnemosyne-localregistered-operator-retrieval | longmemeval-retrievallongmemeval-retrieval-v1 | recall_at_5retrieval | 0.2806 ratioInterval: 0.24880000000000002 to 0.31373333333333336 | View run2026-10-04-longmemeval-retrieval-b0cdbd89operator-run; not publishableOperator: Mnemosyne projectRun disclosuresSystem: mnemosyne-local Track: registered-operator-retrieval Benchmark: longmemeval-retrieval longmemeval-retrieval-v1 Publication: operator-run; not publishable Operator: Mnemosyne project; disclosed: true Metrics
Immutable artifacts
|
| mnemosyne-localregistered-operator-retrieval | longmemeval-retrievallongmemeval-retrieval-v1 | ndcg_at_5retrieval | 0.2967188496001503 ratioInterval: 0.26378166565316336 to 0.33038741454372256 | View run2026-10-04-longmemeval-retrieval-b0cdbd89operator-run; not publishableOperator: Mnemosyne projectRun disclosuresSystem: mnemosyne-local Track: registered-operator-retrieval Benchmark: longmemeval-retrieval longmemeval-retrieval-v1 Publication: operator-run; not publishable Operator: Mnemosyne project; disclosed: true Metrics
Immutable artifacts
|
Retrieval is not answer quality
Recall measures whether useful evidence was found. Answer quality measures whether the response was correct. We report them separately.
Read the methodsFollow the evidence
Every published run links its configuration, uncertainty and question-level traces. Missing measurements stay missing.