Workspace BenchmarksScope data
Official benchmark setup required. Current development checks and adapted results are not official benchmark runs. Official comparisons must use the upstream code, datasets, scoring and prescribed setup, with versions and deviations disclosed. Protocol status →

THE AI MEMORY OBSERVATORY

Beyond recall.
The whole memory.

Explore the benchmarks, inspect the evidence, and understand what it takes to evaluate a complete AI memory system.

Research preview · Operator-run by Mnemosyne

WHOLE-MEMORY SCOPEPlanned evaluation
Memory is a lifecycle.Every stage needs evidence.

Selected dimensions of the proposed suite. Not measured performance.

Coverage is not compatibility. Existing benchmarks can evaluate Mnemosyne. Our current runs measure only part of its memory lifecycle. Understand the limits → · See benchmark-to-test mappings →

THE EVALUATION LANDSCAPE

Find a benchmark

Download catalog ↓

These are scope and implementation labels, not admission badges or performance results. A catalog entry is not a result or proof that its full workload runs on this computer.

14 benchmark families

B01Development components

LongMemEval-S

Retrieval and separate QA adapters exist; the current retrieval registration is a single-system characterization, not an official comparison.

LongMemEval: purpose and limits

Designed to test: Long-conversation question answering: extracting facts, combining sessions, tracking updates, temporal reasoning and abstaining. Retrieval evaluation separately measures whether supporting sessions or turns were found.

Interpretation limit: QA accuracy and retrieval quality answer different questions. Neither alone establishes deletion, authorization or recovery. Our older retrieval scores use different formulas; the new session adapter has only development validation.

Primary source · See testing-element mapping →

Reporting rules

Official protocol and enhanced successor results must remain separate.

B02Planned family

LongMemEval-V2

Planned family; complete official protocol, licensed input pins, adapter and measured evidence remain required.

LongMemEval-V2: purpose and limits

Designed to test: Evaluates static state recall, dynamic state tracking, workflow knowledge, environment gotchas and premise awareness using agent trajectories, including visual context.

Interpretation limit: Trajectory QA is not proof of successful action execution. Image input support, timestamps and trajectory semantics must survive adaptation; a text-only adapter cannot claim full multimodal coverage.

Primary source · See testing-element mapping →

Reporting rules

Official protocol and enhanced successor results must remain separate.

B03Development components

HippoRAG: MuSiQue / 2Wiki / HotpotQA

Retrieval and reader adapter source exists; positive graph-effect and comparable reader evidence remain open.

HippoRAG evaluation datasets: purpose and limits

Designed to test: HippoRAG is a retrieval system and framework. MuSiQue, 2WikiMultiHopQA and HotpotQA supply evidence-retrieval and multi-hop question-answering tasks.

Interpretation limit: A graph-based reference system is not itself a universal memory benchmark. Reader quality and retrieval both affect QA; graph benefit requires a controlled ablation.

Primary source · See testing-element mapping →

Reporting rules

Official protocol and enhanced successor results must remain separate.

B04Development components

MemoryAgentBench

Development adapter source exists; official execution and upstream contribution remain open.

MemoryAgentBench: purpose and limits

Designed to test: Incremental interaction tests accurate retrieval, test-time learning, long-range understanding and conflict resolution.

Interpretation limit: A development adapter does not reproduce the complete upstream task distribution or demonstrate durable storage, access control or operational recovery.

Primary source · See testing-element mapping →

Reporting rules

Official protocol and enhanced successor results must remain separate.

B05Development components

BEAM-1M / BEAM-10M

Development adapter source exists; official scale runs and resource acceptance remain open.

BEAM: purpose and limits

Designed to test: Measures memory abilities over long, coherent conversations, including million-token-scale settings.

Interpretation limit: A small local fixture cannot establish million-token scalability. Context length, ingest time, query latency, model and resource costs must accompany quality scores.

Primary source · See testing-element mapping →

Reporting rules

Official protocol and enhanced successor results must remain separate.

B06Development components

LoCoMo

Ingestion, pinned scoring/tokenizer, native public capture/answer and offline replay components exist. Complete runner integration, dataset license admission, runtime verification and real measured runs remain open.

LoCoMo: purpose and limits

Designed to test: Long-term conversational memory evaluation includes question answering, event summarization and multimodal dialogue generation.

Interpretation limit: Our adapter and replay components do not imply coverage of every upstream task. QA results alone cannot establish summarization or multimodal generation performance.

Primary source · See testing-element mapping →

Reporting rules

Not for leading headline claims under the current publication policy.

B07Planned family

Memora / FAMA

Planned family; complete official protocol, licensed input pins, adapter and measured evidence remain required.

Memora / FAMA: purpose and limits

Designed to test: Memora evaluates conversation-grounded remembering and forgetting. FAMA is a combined evaluation metric, not a second independent benchmark.

Interpretation limit: Not producing deleted information in an answer does not prove its physical removal from stores, indexes or backups. Those require separate residue checks.

Primary source · See testing-element mapping →

Reporting rules

Official protocol and enhanced successor results must remain separate.

B08Planned family

MemoryArena

Planned family; complete official protocol, licensed input pins, adapter and measured evidence remain required.

MemoryArena: purpose and limits

Designed to test: Interdependent, multi-session agent tasks test whether prior actions and feedback help later work, including navigation, planning, search and reasoning.

Interpretation limit: Task success reflects the agent, tools and memory together. Isolating memory contribution requires matched agents and memory ablations.

Primary source · See testing-element mapping →

Reporting rules

Official protocol and enhanced successor results must remain separate.

B09Planned family

AFTER

Planned family; complete official protocol, licensed input pins, adapter and measured evidence remain required.

AFTER: purpose and limits

Designed to test: Evaluates procedural skill revision, specialization, reuse and transfer across tasks, roles and model backbones.

Interpretation limit: Transfer across tested roles does not establish universal transfer or forgetting. This is a procedural-memory evaluation, not a physical-erasure test.

Primary source · See testing-element mapping →

Reporting rules

Official protocol and enhanced successor results must remain separate.

B10Planned family

STATE-Bench

Planned family; complete official protocol, licensed input pins, adapter and measured evidence remain required.

STATE-Bench: purpose and limits

Designed to test: Benchmarks agents on enterprise workflows, including learning through memory, skills and prompt optimization.

Interpretation limit: Enterprise task success is an end-to-end outcome; a matched configuration is needed to attribute gains specifically to memory.

Primary source · See testing-element mapping →

Reporting rules

Official protocol and enhanced successor results must remain separate.

B11Planned family

GroupMemBench / GateMem

Planned family; complete official protocol, licensed input pins, adapter and measured evidence remain required.

GroupMemBench: purpose and limits

Designed to test: Multi-party conversations test group dynamics, speaker-grounded beliefs and audience-specific language. Queries include multi-hop reasoning, updates, ambiguity, user-implicit reasoning, time and abstention.

Interpretation limit: Remembering which person said something differs from enforcing access permissions. Multi-party QA alone is not authorization or tenant-isolation proof.

Primary source · See testing-element mapping →

GateMem: purpose and limits

Designed to test: Shared memory across multiple principals tests useful recall under updates, contextual authorization boundaries and active forgetting after deletion.

Interpretation limit: Answer-level leakage and forgetting checks do not establish physical erasure. Our local signed-identity checks are related engineering checks, not GateMem execution.

Primary source · See testing-element mapping →

Reporting rules

Official protocol and enhanced successor results must remain separate.

B12Development components

PM-Bench / TriggerBench

Repository-specific action probes exist; they are not official upstream results.

PM-Bench: purpose and limits

Designed to test: Prospective-memory evaluation tests remembering delayed intentions and acting on future cues or state changes while other activity continues.

Interpretation limit: Recognizing an intention is not executing it at the correct moment. Local authored action probes do not constitute results on the upstream benchmark.

Primary source · See testing-element mapping →

TriggerBench: purpose and limits

Designed to test: Tests prospective memory in assistant and professional workflows: retaining an intention until its triggering circumstances arrive.

Interpretation limit: A repository-specific probe shares a construct but not necessarily upstream tasks, scoring or difficulty. Official comparability remains unestablished.

Primary source · See testing-element mapping →

Reporting rules

Official protocol and enhanced successor results must remain separate.

B13Planned family

EvoMemBench

Planned family; complete official protocol, licensed input pins, adapter and measured evidence remain required.

EvoMemBench: purpose and limits

Designed to test: Evaluates self-evolving agent memory across within-episode and cross-episode use, spanning knowledge and execution content.

Interpretation limit: Improvement in a particular agent loop does not prove model-independent learning or causal memory benefit without controls.

Primary source · See testing-element mapping →

Reporting rules

Official protocol and enhanced successor results must remain separate.

B14Planned family

EMemBench

Planned family; complete official protocol, licensed input pins, adapter and measured evidence remain required.

EMemBench: purpose and limits

Designed to test: Registered in our slate as an interactive memory benchmark. Detailed task-level mapping is pending source verification.

Interpretation limit: The primary paper could not be fully retrieved in this review. We leave the mapping unassigned rather than infer visual, game or operational coverage from its name.

Primary source · See testing-element mapping →

Reporting rules

Official protocol and enhanced successor results must remain separate.

RAW DEVELOPMENT EVIDENCE

See what actually ran.

Read the failed attempts, controlled experiments and unresolved findings. Download complete records without mistaking development work for admitted results.

Inspect experiments →

BEYOND A SINGLE SCORE

Memory is more than recall.

Correction. Forgetting. Provenance. Future actions. Security. Recovery. Explore the complete planned scope and the evidence still needed.

Explore 20 modules →
What these tests see—and what they miss

Existing benchmarks can evaluate Mnemosyne, but a score does not describe its whole memory system. Coverage means which behaviors a protocol measures. Adapter compatibility means which system interfaces a runner actually exercises. Limited coverage is not evidence of incompatibility.

For Mnemosyne, the accurate limitation is partial measurement of the implemented memory lifecycle, not that conventional benchmarks cannot measure it. A question-answering test can measure the benefit of an internal memory mechanism when that mechanism is enabled during the run. It cannot establish the behavior of operations the run never invokes. Any missing integration belongs to our adapter status, not to an assumption that the benchmark is inherently incompatible.

LongMemEval explicitly supports custom systems: process timestamped histories and submit answers. Its original tasks cover extraction, multi-session reasoning, updates, temporal reasoning and abstention. LoCoMo covers conversational QA, event summarization and multimodal dialog generation; a QA-only run covers only part of that release. These are useful measurements, not tests of every operational property. These statements concern the original protocols, not every newer version or benchmark in this catalog.

Version matters: the LongMemEval repository now points to LongMemEval-V2. The coverage discussion here concerns the original protocol used by our current run; it is not a coverage assessment of V2.

Our current runner has a narrower view

The local LongMemEval retrieval runner batch-captures each question’s history into an isolated store, searches it and scores retrieved session IDs. It measures fractional Recall@5 and conventional nDCG@5, not generated-answer quality. These historical metrics differ from upstream LongMemEval’s all-evidence recall and DCG discount; they must not be presented as unchanged upstream scores. A separate versioned adapter now exercises the upstream formulas on a registered synthetic development suite. A full held-out run under that profile has not been completed. It does not explicitly drive scheduled consolidation or rehearsal over time, branch/merge workflows, recurring actions, deletion verification or tenant-isolation attacks. Those capabilities need their own exercised interfaces and evidence.

For example, answering a question about dates tests temporal reasoning. It does not by itself test whether a memory survives months of interference, whether scheduled rehearsal preserves it, or whether a deletion prevents it from resurfacing. Those are different behaviors requiring explicit workloads and checks. A retrieval score alone cannot establish them.

Mnemosyne already has implementations and regression tests for several of these behaviors, including branch/merge, recurring actions and protected-memory rehearsal. Their complete whole-memory evaluations remain unfinished. Inspect implementation evidence and its limits · See module status.

Match the memory interface and the retrieval budget

A benchmark may divide a document into large chunks while a memory system limits the total context returned by search. A token budget smaller than those chunks can exclude evidence before the answer model sees it. Matching a number called “tokens” is not enough: systems can use different tokenizers or estimates, and memory wrappers also consume space.

Our local MemoryAgentBench integration exposed this exact kind of mismatch with Mnemosyne’s default context budget. That run is a compatibility diagnostic, not a fair basis for ranking answer quality against the official multi-chunk retrieval setup. An expanded-budget experiment must retain the original result, declare the changed settings, and report resource use alongside quality.

The same rule applies to every participating system: verify what was actually stored, what search could return, and what the answer model received. Do not silently shrink the task, raise one system’s budget, drop empty retrievals, or attribute an adapter limitation to the upstream benchmark.

Fair evidence, including weaknesses

A low score remains a real result for the tested configuration; untested features do not cancel it or prove superiority. We must preserve upstream protocols for comparable results and disclose adapter limitations. Additional whole-memory tests belong in a separately identified track, with the same rules available to every participating system.

Three tracks, separate conclusions

Official upstream: the original protocol, inputs and scoring, unchanged. Enhanced successor: separately versioned tests with disclosed differences and controls. Development: local fixtures and conformance tests, never substituted for official results.

LongMemEval-QA is internal-only under the current publication policy. Retrieval and answer quality remain separate. No combined overall rank is implied.

Read the full slate and source audit