Workspace CoverageScope data
Official benchmark setup required. Current development checks and adapted results are not official benchmark runs. Official comparisons must use the upstream code, datasets, scoring and prescribed setup, with versions and deviations disclosed. Protocol status →

What complete memory needs to prove

24 capabilities. 20 modules. Evidence for how they work together.

The Mnemetric Whole-Memory Benchmark is our own suite within this multi-benchmark platform. This is the scope of the benchmark program, not a measured system score. It does not certify Mnemosyne or any other system. Per-system support and quality need separately verified adapters and runs.

Reviewed 2026-10-04. Read the whole-memory specification · Inspect the implementation audit

Our coverage target

Mnemetric aims to test the union of capabilities measured by leading memory benchmarks and the full observable capabilities of Mnemosyne, including behavior that existing tests leave unmeasured. The current catalog is a starting inventory, not a ceiling or a claim of complete coverage.

Official benchmark results stay separate from adapted tests. Custom lifecycle tests use neutral behavior contracts so other memory systems can participate. Every capability needs a concrete test, a scoring method and retained evidence; missing adapters, unsupported features and unmeasured results remain visible.

Capability map

CapabilityBenchmark modules
C01 · Capture and normalizationM01, M02
C02 · Durable persistence and replayable projectionsM01, M15, M17
C03 · Semantic, lexical, graph, and temporal retrievalM02
C04 · Temporal evolution and bitemporal as-of queriesM03
C05 · Correction, contradiction, multi-hypothesis belief, branch/mergeM04
C06 · Provenance, citation, explanation, and audit lineageM05, M20
C07 · Entities, relations, preferences, user models, and proceduresM02, M14
C08 · Consolidation, reflection, and test-time learningM06, M14
C09 · Retention, rehearsal, decay, fidelity demotion, and pruningM07
C10 · Reversible forgetting and tombstoningM08
C11 · Declared-surface irreversible erasure and residue controlM09
C12 · Calibrated uncertainty and abstentionM10
C13 · Prospective memory and deferred actionM12
C14 · Working and session memoryM13
C15 · Retrieved-content safety, trust, and write authorityM11
C16 · Authentication, authorization, tenant/user/group isolationM11
C17 · Determinism, reproducibility, and replayM15, M20
C18 · Backend parity and portabilityM16
C19 · Operational custody, queues, retries, backup/restore, and recoveryM17
C20 · Agent-facing ABI, MCP/CLI transports, and interchangeM16, M18
C21 · Multimodal evidence and custodyM19
C22 · Public result contract, ledger, rendering, and publicationM20
C23 · End-to-end task and procedural utilityM14
C24 · Efficiency, scale, and resource behaviorM01, M02, M03, M04, M05, M06, M07, M08, M09, M10, M11, M12, M13, M14, M15, M16, M17, M18, M19

Module implementation and remaining evidence

M01 · Capture/durability

Local development pilot exists; broader backend and failure/durability evidence remains.

M02 · Retrieval/organization

Registered development cell and replay exist; full external/comparative acceptance remains.

M03 · Temporal evolution

Valid-time slice exists; full transaction-time bitemporality remains.

M04 · Conflict/correction

Development matrix and replay exist; full public acceptance remains.

M05 · Provenance/explanation

Development cell exists; quarantines and complete outcome/lineage evidence remain.

M06 · Consolidation/learning

Stage A/B development cell exists; sustained learning and non-degradation acceptance remains.

M07 · Retention/rehearsal/decay

Public CLI rehearsal scheduling tested across five calendar starts spanning over a virtual year; full seeded retention, retrieval, interference, storage/cost and resource admission remain open.

M08 · Reversible forgetting

Public tombstoning exists; its reversible-delete adapter contract is unresolved. Dedicated fixture, scorer and integration remain missing; restore correctness is not claimed.

M09 · Declared-surface erasure

Plan exists; full benchmark and explicit surface/restore evidence missing.

M10 · Calibration/abstention

Deterministic pilot exists; model-backed calibration and public-label evidence remain.

M11 · Security/isolation

Plan and separate security-development infrastructure exist; full M11 benchmark missing.

M12 · Prospective action

Development captures include 130 five-trigger receipts, recovery with 50 injected response losses, and 1,220 bounded fan-out operations with 750 receipts and exact replay. A separate resource run sampled at most 88 MiB RSS; this is not resource admission. The 220-conversation corpus has a public formation bridge, state/timing diagnostics and saved-trace verification. Local Qwen3 0.6B attempts remain incomplete. The schema variant scheduled no tasks and missed 17 eligible occurrences across 21 completed cases before an HTTP error; this is not a full-corpus score. A single-input follow-up exposed schema-order sensitivity, not correct formation. Draft reference captures remain unadmitted, with a verified nested-condition policy difference. Live PostgreSQL retry/BurnOS checks passed at 55ae2ef9. Full overload/implicit execution, reference calibration, cost and admission remain open.

M13 · Working memory

Development probe exists; capacity and promotion-control experiment remains.

M14 · Procedural task utility

Plan exists; closed-agent environment, policy/baseline/inference contract and implementation remain.

M15 · Determinism/replay

Pilot composition exists; all admitted payloads and real clean-checkout reproduction remain.

M16 · Backend/transport parity

Bounded local cassette exists; broader pairs, migration, resource and acceptance evidence remain.

M17 · Custody/recovery

Plan exists; benchmark implementation and declared fault/recovery scope remain.

M18 · Interoperability

Plan exists; benchmark implementation and fair cross-system interchange contracts remain.

M19 · Multimodal memory

Plan exists, module deferred; portable modality contracts, rights, resources and implementation remain.

M20 · Publication integrity

Result-v1/v2, signed ledger, renderer and publication/readiness infrastructure exist; operational admission and full public evidence remain.

Test the interactions too

These joint scenarios remain planned until the relevant modules and their combined behavior have measured evidence. Passing an isolated module does not establish the joint result.

Explore the benchmark catalog · Download scope catalog (JSON)

What maps to what?

Benchmark task → memory behavior → adapter → development check.

These are our source-backed interpretations of overlapping constructs, not claims of equivalent protocols or difficulty. A benchmark may do its intended job well without testing every stage of a memory lifecycle. Missing integration here is our implementation gap, not evidence that the benchmark is incompatible with Mnemosyne.

The same rules apply to every memory system: preserve the task, data, timing, model, budget and scorer; disclose adaptation; report unsupported operations separately from failures. A system profile is not an implemented adapter. These development integrations do not establish competitor support or rankings.

Family-level construct map plus selected reviewed development checks. Not an exhaustive per-task or per-test audit. Reviewed 2026-10-04. Download mapping JSON.

BenchmarkTesting elements → modules / capabilitiesAdapter / system under testLimit of this mapping
LongMemEval
  • Session evidence retrieval → M02 (C03)
  • Temporal questions → M03 (C04)
  • Knowledge updates → M04 (C05)
  • Abstention → M10 (C12)
eval/public/adapters/longmemeval_session.py
tests/test_longmemeval_session_adapter.py

Mnemosyne development integration

QA accuracy and retrieval quality answer different questions. Neither alone establishes deletion, authorization or recovery. Our older retrieval scores use different formulas; the new session adapter has only development validation. Source
LongMemEval-V2
  • State recall and change → M02, M03, M04 (C03, C04, C05)
  • Workflow and environment knowledge → M14 (C07, C23)
  • Premise awareness → M10 (C12)
  • Visual trajectory context → M19 (C21)
Integration not claimed
Reviewed check link pending

No implemented system integration claimed

Trajectory QA is not proof of successful action execution. Image input support, timestamps and trajectory semantics must survive adaptation; a text-only adapter cannot claim full multimodal coverage. Source
HippoRAG evaluation datasets
  • Multi-hop evidence retrieval and answering → M02, M14 (C03, C23)
eval/public/adapters/hipporag_multihop.py
tests/test_public_hipporag.py

Mnemosyne development integration

A graph-based reference system is not itself a universal memory benchmark. Reader quality and retrieval both affect QA; graph benefit requires a controlled ablation. Source
MemoryAgentBench
  • Retrieval and long-range understanding → M02 (C03)
  • Learning from experience → M06, M14 (C08, C23)
  • Conflict resolution → M04 (C05)
eval/public/adapters/memoryagentbench.py
tests/test_public_memoryagentbench.py

Mnemosyne development integration

A development adapter does not reproduce the complete upstream task distribution or demonstrate durable storage, access control or operational recovery. Source
BEAM
  • Long-conversation memory → M02, M03, M04 (C03, C04, C05)
  • Scale and resource accounting → M01, M02 (C24)
eval/public/adapters/beam.py
tests/test_public_beam.py

Mnemosyne development integration

A small local fixture cannot establish million-token scalability. Context length, ingest time, query latency, model and resource costs must accompany quality scores. Source
LoCoMo
  • Conversational QA and temporal events → M02, M03 (C03, C04)
  • Event summarization → M06 (C08)
  • Multimodal dialogue → M19 (C21)
eval/public/adapters/locomo.py
Reviewed check link pending

Mnemosyne development integration

Our adapter and replay components do not imply coverage of every upstream task. QA results alone cannot establish summarization or multimodal generation performance. Source
Memora / FAMA
  • Remembering and reasoning → M02 (C03)
  • Updated or deleted information in answers → M04, M08 (C05, C10)
Integration not claimed
Reviewed check link pending

No implemented system integration claimed

Not producing deleted information in an answer does not prove its physical removal from stores, indexes or backups. Those require separate residue checks. Source
MemoryArena
  • Experience reuse in multi-session action → M06, M14 (C08, C23)
Integration not claimed
Reviewed check link pending

No implemented system integration claimed

Task success reflects the agent, tools and memory together. Isolating memory contribution requires matched agents and memory ablations. Source
AFTER
  • Procedural learning and transfer → M06, M14 (C07, C08, C23)
Integration not claimed
Reviewed check link pending

No implemented system integration claimed

Transfer across tested roles does not establish universal transfer or forgetting. This is a procedural-memory evaluation, not a physical-erasure test. Source
STATE-Bench
  • Workflow execution and learning → M06, M14 (C08, C23)
Integration not claimed
Reviewed check link pending

No implemented system integration claimed

Enterprise task success is an end-to-end outcome; a matched configuration is needed to attribute gains specifically to memory. Source
GroupMemBench
  • Speaker and audience-grounded retrieval → M02 (C03, C07)
  • Updates and temporal reasoning → M03, M04 (C04, C05)
  • Abstention → M10 (C12)
Integration not claimed
Reviewed check link pending

No implemented system integration claimed

Remembering which person said something differs from enforcing access permissions. Multi-party QA alone is not authorization or tenant-isolation proof. Source
GateMem
  • Authorization-aware recall → M02, M11 (C03, C16)
  • Updates and active forgetting → M04, M08 (C05, C10)
Integration not claimed
Reviewed check link pending

No implemented system integration claimed

Answer-level leakage and forgetting checks do not establish physical erasure. Our local signed-identity checks are related engineering checks, not GateMem execution. Source
PM-Bench
  • Delayed intentions and triggering → M12 (C13)
eval/public/adapters/pm_bench_triggerbench.py
tests/test_public_pm_bench_triggerbench.py

Mnemosyne development integration

Recognizing an intention is not executing it at the correct moment. Local authored action probes do not constitute results on the upstream benchmark. Source
TriggerBench
  • Cue-dependent future action → M12 (C13)
eval/public/adapters/pm_bench_triggerbench.py
tests/test_public_pm_bench_triggerbench.py

Mnemosyne development integration

A repository-specific probe shares a construct but not necessarily upstream tasks, scoring or difficulty. Official comparability remains unestablished. Source
EvoMemBench
  • Knowledge and execution experience reuse → M06, M14 (C08, C23)
Integration not claimed
Reviewed check link pending

No implemented system integration claimed

Improvement in a particular agent loop does not prove model-independent learning or causal memory benefit without controls. Source
EMemBench
    Task mapping pending
    Integration not claimed
    Reviewed check link pending

    No implemented system integration claimed

    The primary paper could not be fully retrieved in this review. We leave the mapping unassigned rather than infer visual, game or operational coverage from its name. Source

    Development checks and their limits

    Selected checks below test integration or evaluator behavior. They are not official benchmark scores. The working-memory audit below provides more detailed positive checks and unresolved cases.

    Check / source fileRelated benchmark / moduleWhat it checksWhat remains unproven
    Session metric formulas
    tests/test_longmemeval_upstream_metrics.py
    B01
    M02
    Matches audited upstream session discount and recall definitions at 5 and 10; checks bundle recomputation.Full dataset, turn-level metrics and official QA judge remain separate.
    Session adapter protocol
    tests/test_longmemeval_session_adapter.py
    B01
    M02
    Synthetic normalization, explicit retrieval abstention exclusion, ten-hit retention, empty retrieval, date-prefix transport and real CLI bundle.Excluding abstention questions from retrieval does not test QA abstention. Preserving date text does not establish temporal reasoning, bitemporal storage or full official execution.
    Multi-hop adapter
    tests/test_public_hipporag.py
    B03
    M02, M14
    Development adapter contract for multi-hop retrieval and answering.Not a controlled comparison with HippoRAG or proof of graph benefit.
    Incremental memory adapter
    tests/test_public_memoryagentbench.py
    B04
    M02, M04, M06
    Development adapter contract.Full upstream categories and official scoring need protocol audit and retained runs.
    Long-context adapter
    tests/test_public_beam.py
    B05
    M02
    Development adapter contract.Small fixtures do not establish 1M/10M-scale performance.
    Prospective action probes
    tests/test_public_pm_bench_triggerbench.py
    B12
    M12
    Authored development cases related to future-action constructs.Not official PM-Bench or TriggerBench tasks or scores.
    Working-memory selection and abstention
    tests/test_public_working_memory_action_probe.py
    B01
    M10, M13
    Related construct only: explicit relevance and deterministic policy checks.Not semantic QA abstention, learned relevance or full working-memory capacity.
    Scope identity and TTL expiry
    tests/test_working_memory_cli_contract.py
    B11
    M11, M13
    GateMem-related authorization boundary checks; TTL and scoped expiry are independent engineering contracts.Not GroupMemBench group reasoning or GateMem workload execution. TTL is not temporal-question reasoning.
    CLI/MCP declarations
    tests/test_cli_coverage_inventory.py
    No direct upstream equivalent assigned
    M16, M18
    Static interface inventory helps locate integration surfaces.Command counts are not behavior coverage; MCP inventory is separately available below.
    MCP declarations
    tests/test_mcp_coverage_inventory.py
    No direct upstream equivalent assigned
    M16, M18
    Static tool inventory and signature correspondence.Does not establish transport parity or successful tool execution.
    Backend/transport parity
    tests/test_public_wmbs_m16.py
    No direct upstream equivalent assigned
    M16
    Bounded local parity contract; no direct upstream task equivalence claimed.Full backend, SDK and remote transport matrix remains open.
    Result publication contracts
    tests/test_leaderboard_render.py
    No direct upstream equivalent assigned
    M20
    Checks evidence presentation and publication behavior.Rendering checks say nothing about memory-system quality.

    From capability claims to test evidence

    120 CLI commands and 60 MCP tools inventoried. Behavioral coverage remains incomplete.

    These interfaces overlap; the counts are not distinct memory capabilities or passed tests. We are auditing each behavior against executable tests and retained evidence.

    Runtime revision: e1ad2d0cf5ac. Reviewed 2026-10-04.

    First behavior audit: working memory

    The checks below are bounded local development tests through the real CLI, plus query-filter checks through local MCP stdio. They are not published benchmark scores, comparisons against other systems, or full M13 acceptance. MCP HTTP/SDK transport and full lifecycle parity remain unverified. Relevance is supplied explicitly in item text; the action checks do not establish learned semantic relevance.

    InterfaceReviewed development checksStill needs evidence
    working-seedMCP: working_seed
    • positive action selection using explicit relevance
    • capacity/overflow
    • concurrent writes
    • invalid evidence
    • trust policy
    • durability after forced termination
    working-queryMCP: working_query
    • positive action selection
    • no visibility at exact TTL boundary
    • below-threshold policy abstention with item still visible
    • separate tenant/session/user/agent/task/branch scopes in a shared store
    • signed identity mismatch rejected for tenant/user/agent/session
    • MCP stdio tool discovery advertises limit and kinds
    • MCP newest-first ordering, limit, kind filter and combined filter-before-limit
    • MCP negative and boolean limits rejected by schema
    • concurrent and malformed-input isolation stress
    • retrieval relevance without explicit numeric relevance labels
    • MCP HTTP/SDK transport parity and authorization adversarial matrix
    working-expireMCP: working_expire
    • due item physically removed from selected scope
    • repeated expiry preserves live item
    • adjacent tenant/session/user/agent/task/branch scopes preserved
    • concurrent expiry/write races
    • crash recovery during expiry
    • MCP transport equivalence
    working-promoteMCP: working_promoteNo reviewed execution evidence here.
    • successful and rejected promotions
    • evidence retention
    • promotion utility
    • atomicity and retry

    CLI inventory · MCP inventory · Behavior review

    Upstream task coverage audit

    Two benchmark versions have an initial, source-pinned task inventory. These are documentation audit rows, not executed benchmark tests. Scorer review, adapter mapping and result evidence remain open; the rest of the benchmark catalog still needs this task-level audit. The original benchmark’s abstention row is a cross-cutting classification, not a seventh mutually exclusive question type.

    LongMemEval: source audit started

    single-session-user, single-session-assistant, single-session-preference, temporal-reasoning, knowledge-update, multi-session, abstention.

    • Keep S, M and oracle variants separate.
    • Abstention is identified by the _abs question-ID suffix; retrieval evaluation excludes those instances.
    • QA correctness and turn/session retrieval metrics are different outcomes.
    • Upstream recall_all requires every gold session; current historical Mnemetric recall_at_5 is fractional.
    • Upstream DCG weights its first two ranks equally; historical Mnemetric nDCG uses the conventional discount.
    • QA aggregation includes per-type, macro, overall and abstention accuracy; the script asserts judge gpt-4o-2024-08-06.

    Pinned upstream source

    LongMemEval-V2: source audit started

    Static state recall, Dynamic state tracking, Workflow knowledge, Environment gotchas, Premise awareness.

    • Keep web/enterprise domains and small/medium tiers identifiable.
    • Memory accepts trajectories and returns text/image context under a token budget.
    • Query receives question text and optional image, not gold labels or benchmark metadata.
    • Accuracy, query latency and LAFS frontier gain require separate reporting.

    Pinned upstream source

    Download upstream task audit (JSON)