What complete memory needs to prove
24 capabilities. 20 modules. Evidence for how they work together.
The Mnemetric Whole-Memory Benchmark is our own suite within this multi-benchmark platform. This is the scope of the benchmark program, not a measured system score. It does not certify Mnemosyne or any other system. Per-system support and quality need separately verified adapters and runs.
Reviewed 2026-10-04. Read the whole-memory specification · Inspect the implementation audit
Our coverage target
Mnemetric aims to test the union of capabilities measured by leading memory benchmarks and the full observable capabilities of Mnemosyne, including behavior that existing tests leave unmeasured. The current catalog is a starting inventory, not a ceiling or a claim of complete coverage.
Official benchmark results stay separate from adapted tests. Custom lifecycle tests use neutral behavior contracts so other memory systems can participate. Every capability needs a concrete test, a scoring method and retained evidence; missing adapters, unsupported features and unmeasured results remain visible.
Capability map
| Capability | Benchmark modules |
|---|---|
| C01 · Capture and normalization | M01, M02 |
| C02 · Durable persistence and replayable projections | M01, M15, M17 |
| C03 · Semantic, lexical, graph, and temporal retrieval | M02 |
| C04 · Temporal evolution and bitemporal as-of queries | M03 |
| C05 · Correction, contradiction, multi-hypothesis belief, branch/merge | M04 |
| C06 · Provenance, citation, explanation, and audit lineage | M05, M20 |
| C07 · Entities, relations, preferences, user models, and procedures | M02, M14 |
| C08 · Consolidation, reflection, and test-time learning | M06, M14 |
| C09 · Retention, rehearsal, decay, fidelity demotion, and pruning | M07 |
| C10 · Reversible forgetting and tombstoning | M08 |
| C11 · Declared-surface irreversible erasure and residue control | M09 |
| C12 · Calibrated uncertainty and abstention | M10 |
| C13 · Prospective memory and deferred action | M12 |
| C14 · Working and session memory | M13 |
| C15 · Retrieved-content safety, trust, and write authority | M11 |
| C16 · Authentication, authorization, tenant/user/group isolation | M11 |
| C17 · Determinism, reproducibility, and replay | M15, M20 |
| C18 · Backend parity and portability | M16 |
| C19 · Operational custody, queues, retries, backup/restore, and recovery | M17 |
| C20 · Agent-facing ABI, MCP/CLI transports, and interchange | M16, M18 |
| C21 · Multimodal evidence and custody | M19 |
| C22 · Public result contract, ledger, rendering, and publication | M20 |
| C23 · End-to-end task and procedural utility | M14 |
| C24 · Efficiency, scale, and resource behavior | M01, M02, M03, M04, M05, M06, M07, M08, M09, M10, M11, M12, M13, M14, M15, M16, M17, M18, M19 |
Module implementation and remaining evidence
M01 · Capture/durability
Local development pilot exists; broader backend and failure/durability evidence remains.
M02 · Retrieval/organization
Registered development cell and replay exist; full external/comparative acceptance remains.
M03 · Temporal evolution
Valid-time slice exists; full transaction-time bitemporality remains.
M04 · Conflict/correction
Development matrix and replay exist; full public acceptance remains.
M05 · Provenance/explanation
Development cell exists; quarantines and complete outcome/lineage evidence remain.
M06 · Consolidation/learning
Stage A/B development cell exists; sustained learning and non-degradation acceptance remains.
M07 · Retention/rehearsal/decay
Public CLI rehearsal scheduling tested across five calendar starts spanning over a virtual year; full seeded retention, retrieval, interference, storage/cost and resource admission remain open.
M08 · Reversible forgetting
Public tombstoning exists; its reversible-delete adapter contract is unresolved. Dedicated fixture, scorer and integration remain missing; restore correctness is not claimed.
M09 · Declared-surface erasure
Plan exists; full benchmark and explicit surface/restore evidence missing.
M10 · Calibration/abstention
Deterministic pilot exists; model-backed calibration and public-label evidence remain.
M11 · Security/isolation
Plan and separate security-development infrastructure exist; full M11 benchmark missing.
M12 · Prospective action
Development captures include 130 five-trigger receipts, recovery with 50 injected response losses, and 1,220 bounded fan-out operations with 750 receipts and exact replay. A separate resource run sampled at most 88 MiB RSS; this is not resource admission. The 220-conversation corpus has a public formation bridge, state/timing diagnostics and saved-trace verification. Local Qwen3 0.6B attempts remain incomplete. The schema variant scheduled no tasks and missed 17 eligible occurrences across 21 completed cases before an HTTP error; this is not a full-corpus score. A single-input follow-up exposed schema-order sensitivity, not correct formation. Draft reference captures remain unadmitted, with a verified nested-condition policy difference. Live PostgreSQL retry/BurnOS checks passed at 55ae2ef9. Full overload/implicit execution, reference calibration, cost and admission remain open.
M13 · Working memory
Development probe exists; capacity and promotion-control experiment remains.
M14 · Procedural task utility
Plan exists; closed-agent environment, policy/baseline/inference contract and implementation remain.
M15 · Determinism/replay
Pilot composition exists; all admitted payloads and real clean-checkout reproduction remain.
M16 · Backend/transport parity
Bounded local cassette exists; broader pairs, migration, resource and acceptance evidence remain.
M17 · Custody/recovery
Plan exists; benchmark implementation and declared fault/recovery scope remain.
M18 · Interoperability
Plan exists; benchmark implementation and fair cross-system interchange contracts remain.
M19 · Multimodal memory
Plan exists, module deferred; portable modality contracts, rights, resources and implementation remain.
M20 · Publication integrity
Result-v1/v2, signed ledger, renderer and publication/readiness infrastructure exist; operational admission and full public evidence remain.
Test the interactions too
- Correction with evidence: M03, M04, M05
- Retain, forget, erase: M07, M08, M09
- Learn without leaking: M06, M11, M14
- Authorized future action: M11, M12, M17
- Working-to-long-term promotion: M10, M13, M16
- Recoverable public evidence: M15, M16, M17, M20
These joint scenarios remain planned until the relevant modules and their combined behavior have measured evidence. Passing an isolated module does not establish the joint result.
Explore the benchmark catalog · Download scope catalog (JSON)
What maps to what?
Benchmark task → memory behavior → adapter → development check.
These are our source-backed interpretations of overlapping constructs, not claims of equivalent protocols or difficulty. A benchmark may do its intended job well without testing every stage of a memory lifecycle. Missing integration here is our implementation gap, not evidence that the benchmark is incompatible with Mnemosyne.
The same rules apply to every memory system: preserve the task, data, timing, model, budget and scorer; disclose adaptation; report unsupported operations separately from failures. A system profile is not an implemented adapter. These development integrations do not establish competitor support or rankings.
Family-level construct map plus selected reviewed development checks. Not an exhaustive per-task or per-test audit. Reviewed 2026-10-04. Download mapping JSON.
| Benchmark | Testing elements → modules / capabilities | Adapter / system under test | Limit of this mapping |
|---|---|---|---|
| LongMemEval | eval/public/adapters/longmemeval_session.pytests/test_longmemeval_session_adapter.pyMnemosyne development integration | QA accuracy and retrieval quality answer different questions. Neither alone establishes deletion, authorization or recovery. Our older retrieval scores use different formulas; the new session adapter has only development validation. Source | |
| LongMemEval-V2 | Integration not claimed Reviewed check link pending No implemented system integration claimed | Trajectory QA is not proof of successful action execution. Image input support, timestamps and trajectory semantics must survive adaptation; a text-only adapter cannot claim full multimodal coverage. Source | |
| HippoRAG evaluation datasets | eval/public/adapters/hipporag_multihop.pytests/test_public_hipporag.pyMnemosyne development integration | A graph-based reference system is not itself a universal memory benchmark. Reader quality and retrieval both affect QA; graph benefit requires a controlled ablation. Source | |
| MemoryAgentBench | eval/public/adapters/memoryagentbench.pytests/test_public_memoryagentbench.pyMnemosyne development integration | A development adapter does not reproduce the complete upstream task distribution or demonstrate durable storage, access control or operational recovery. Source | |
| BEAM | eval/public/adapters/beam.pytests/test_public_beam.pyMnemosyne development integration | A small local fixture cannot establish million-token scalability. Context length, ingest time, query latency, model and resource costs must accompany quality scores. Source | |
| LoCoMo | eval/public/adapters/locomo.pyReviewed check link pending Mnemosyne development integration | Our adapter and replay components do not imply coverage of every upstream task. QA results alone cannot establish summarization or multimodal generation performance. Source | |
| Memora / FAMA | Integration not claimed Reviewed check link pending No implemented system integration claimed | Not producing deleted information in an answer does not prove its physical removal from stores, indexes or backups. Those require separate residue checks. Source | |
| MemoryArena | Integration not claimed Reviewed check link pending No implemented system integration claimed | Task success reflects the agent, tools and memory together. Isolating memory contribution requires matched agents and memory ablations. Source | |
| AFTER | Integration not claimed Reviewed check link pending No implemented system integration claimed | Transfer across tested roles does not establish universal transfer or forgetting. This is a procedural-memory evaluation, not a physical-erasure test. Source | |
| STATE-Bench | Integration not claimed Reviewed check link pending No implemented system integration claimed | Enterprise task success is an end-to-end outcome; a matched configuration is needed to attribute gains specifically to memory. Source | |
| GroupMemBench | Integration not claimed Reviewed check link pending No implemented system integration claimed | Remembering which person said something differs from enforcing access permissions. Multi-party QA alone is not authorization or tenant-isolation proof. Source | |
| GateMem | Integration not claimed Reviewed check link pending No implemented system integration claimed | Answer-level leakage and forgetting checks do not establish physical erasure. Our local signed-identity checks are related engineering checks, not GateMem execution. Source | |
| PM-Bench |
| eval/public/adapters/pm_bench_triggerbench.pytests/test_public_pm_bench_triggerbench.pyMnemosyne development integration | Recognizing an intention is not executing it at the correct moment. Local authored action probes do not constitute results on the upstream benchmark. Source |
| TriggerBench |
| eval/public/adapters/pm_bench_triggerbench.pytests/test_public_pm_bench_triggerbench.pyMnemosyne development integration | A repository-specific probe shares a construct but not necessarily upstream tasks, scoring or difficulty. Official comparability remains unestablished. Source |
| EvoMemBench | Integration not claimed Reviewed check link pending No implemented system integration claimed | Improvement in a particular agent loop does not prove model-independent learning or causal memory benefit without controls. Source | |
| EMemBench | Integration not claimed Reviewed check link pending No implemented system integration claimed | The primary paper could not be fully retrieved in this review. We leave the mapping unassigned rather than infer visual, game or operational coverage from its name. Source |
Development checks and their limits
Selected checks below test integration or evaluator behavior. They are not official benchmark scores. The working-memory audit below provides more detailed positive checks and unresolved cases.
| Check / source file | Related benchmark / module | What it checks | What remains unproven |
|---|---|---|---|
Session metric formulastests/test_longmemeval_upstream_metrics.py | B01 M02 | Matches audited upstream session discount and recall definitions at 5 and 10; checks bundle recomputation. | Full dataset, turn-level metrics and official QA judge remain separate. |
Session adapter protocoltests/test_longmemeval_session_adapter.py | B01 M02 | Synthetic normalization, explicit retrieval abstention exclusion, ten-hit retention, empty retrieval, date-prefix transport and real CLI bundle. | Excluding abstention questions from retrieval does not test QA abstention. Preserving date text does not establish temporal reasoning, bitemporal storage or full official execution. |
Multi-hop adaptertests/test_public_hipporag.py | B03 M02, M14 | Development adapter contract for multi-hop retrieval and answering. | Not a controlled comparison with HippoRAG or proof of graph benefit. |
Incremental memory adaptertests/test_public_memoryagentbench.py | B04 M02, M04, M06 | Development adapter contract. | Full upstream categories and official scoring need protocol audit and retained runs. |
Long-context adaptertests/test_public_beam.py | B05 M02 | Development adapter contract. | Small fixtures do not establish 1M/10M-scale performance. |
Prospective action probestests/test_public_pm_bench_triggerbench.py | B12 M12 | Authored development cases related to future-action constructs. | Not official PM-Bench or TriggerBench tasks or scores. |
Working-memory selection and abstentiontests/test_public_working_memory_action_probe.py | B01 M10, M13 | Related construct only: explicit relevance and deterministic policy checks. | Not semantic QA abstention, learned relevance or full working-memory capacity. |
Scope identity and TTL expirytests/test_working_memory_cli_contract.py | B11 M11, M13 | GateMem-related authorization boundary checks; TTL and scoped expiry are independent engineering contracts. | Not GroupMemBench group reasoning or GateMem workload execution. TTL is not temporal-question reasoning. |
CLI/MCP declarationstests/test_cli_coverage_inventory.py | No direct upstream equivalent assigned M16, M18 | Static interface inventory helps locate integration surfaces. | Command counts are not behavior coverage; MCP inventory is separately available below. |
MCP declarationstests/test_mcp_coverage_inventory.py | No direct upstream equivalent assigned M16, M18 | Static tool inventory and signature correspondence. | Does not establish transport parity or successful tool execution. |
Backend/transport paritytests/test_public_wmbs_m16.py | No direct upstream equivalent assigned M16 | Bounded local parity contract; no direct upstream task equivalence claimed. | Full backend, SDK and remote transport matrix remains open. |
Result publication contractstests/test_leaderboard_render.py | No direct upstream equivalent assigned M20 | Checks evidence presentation and publication behavior. | Rendering checks say nothing about memory-system quality. |
From capability claims to test evidence
120 CLI commands and 60 MCP tools inventoried. Behavioral coverage remains incomplete.
These interfaces overlap; the counts are not distinct memory capabilities or passed tests. We are auditing each behavior against executable tests and retained evidence.
Runtime revision: e1ad2d0cf5ac. Reviewed 2026-10-04.
First behavior audit: working memory
The checks below are bounded local development tests through the real CLI, plus query-filter checks through local MCP stdio. They are not published benchmark scores, comparisons against other systems, or full M13 acceptance. MCP HTTP/SDK transport and full lifecycle parity remain unverified. Relevance is supplied explicitly in item text; the action checks do not establish learned semantic relevance.
| Interface | Reviewed development checks | Still needs evidence |
|---|---|---|
working-seedMCP: working_seed |
|
|
working-queryMCP: working_query |
|
|
working-expireMCP: working_expire |
|
|
working-promoteMCP: working_promote | No reviewed execution evidence here. |
|
Upstream task coverage audit
Two benchmark versions have an initial, source-pinned task inventory. These are documentation audit rows, not executed benchmark tests. Scorer review, adapter mapping and result evidence remain open; the rest of the benchmark catalog still needs this task-level audit. The original benchmark’s abstention row is a cross-cutting classification, not a seventh mutually exclusive question type.
LongMemEval: source audit started
single-session-user, single-session-assistant, single-session-preference, temporal-reasoning, knowledge-update, multi-session, abstention.
- Keep S, M and oracle variants separate.
- Abstention is identified by the _abs question-ID suffix; retrieval evaluation excludes those instances.
- QA correctness and turn/session retrieval metrics are different outcomes.
- Upstream recall_all requires every gold session; current historical Mnemetric recall_at_5 is fractional.
- Upstream DCG weights its first two ranks equally; historical Mnemetric nDCG uses the conventional discount.
- QA aggregation includes per-type, macro, overall and abstention accuracy; the script asserts judge gpt-4o-2024-08-06.
LongMemEval-V2: source audit started
Static state recall, Dynamic state tracking, Workflow knowledge, Environment gotchas, Premise awareness.
- Keep web/enterprise domains and small/medium tiers identifiable.
- Memory accepts trajectories and returns text/image context under a token budget.
- Query receives question text and optional image, not gold labels or benchmark metadata.
- Accuracy, query latency and LAFS frontier gain require separate reporting.