Why memory needs more than a recall score
The problem is not a shortage of useful tests. It is the distance between a good benchmark score and evidence that a memory system will behave reliably throughout its life.
Mnemetric is being built to close that distance: bring existing benchmarks, their limitations, system capabilities and inspectable results into one place. Its proposed Whole-Memory Benchmark adds explicit tests for the wider lifecycle. This is a coverage and evidence goal, not a claim that the complete suite has already been delivered.
Our core definition
An external AI memory system maintains persistent information outside a model's weights and active context, and provides mechanisms to store, retrieve and manage that information for use across separate interactions or sessions.
This is Mnemetric's operational definition for selecting comparison subjects, not a claim of an industry-wide standard or human-like memory.
- Persistence: stored information survives the end of the interaction or active context.
- Reuse: a model or agent can access that information in a later interaction.
- Management: the evaluated system provides a write path and operations for maintaining its stored information. Manual and automatic writes both qualify, but must be disclosed.
To demonstrate these criteria, store a new test fact, start a fresh session without replaying the original conversation, and retrieve the fact through the memory interface. Identify the external store, write path and retrieval path. This establishes basic eligibility; it does not establish quality, durability under failure or complete lifecycle support.
A vector database alone is a component. The evaluated system includes the mechanisms that write and retrieve its contents for the agent. A file-backed memory can qualify under the same criteria. Fixed model knowledge, a larger context window alone and conversation replay without a persistent managed memory layer are comparison baselines.
Corrections, deletion, automatic learning, temporal reasoning, provenance and permissions are dimensions to measure, not mandatory admission requirements that exclude weaker systems. Unsupported, failed and unmeasured capabilities remain visible. Cross-session persistence does not imply crash recovery, physical erasure or cross-user isolation.
Our comparison population
We evaluate persistent memory layers that models and agents use across interactions: systems such as Mnemosyne, Mem0, GBrain and MemPalace, alongside Graphiti/Zep, Letta, Cognee, MemOS and Supermemory. The system profiles identify the projects and sources. This is a comparison candidate set, not a proven ranking of the best systems.
A larger context window or knowledge stored in model weights is not the main product category here. Long-context models and no-memory configurations provide useful baselines. External memory often places retrieved information into the model's context; that delivery mechanism does not change its status as persistent external memory.
The question is how well the memory layer captures, retains, updates, retrieves and governs information over time. To isolate its contribution, comparisons should control the model and resource budget where feasible and disclose native agent configurations separately.
The decision a score cannot make for you
When choosing AI memory, you need to know more than whether it can find an old sentence. Can it recognize that the sentence is outdated? Explain where its answer came from? Keep one person's information away from another? Recover after a failure? Honor a deletion request? Meet those requirements within a realistic time and cost budget?
A benchmark answers the questions its tasks and scoring rules actually ask. A high score cannot establish an unmeasured property. This is why a strong result on one benchmark may be insufficient for a deployment decision, even when the benchmark is well designed and the result is genuine.
One ordinary week, several different tests
This fictional example illustrates the distinction; it is not a measured result for any system.
- Monday: capture. You tell an assistant your delivery address. It must store the right information with its source.
- Tuesday: correction. You move. It must use the new address for today's delivery while preserving the meaning of an earlier order.
- Wednesday: access. A colleague asks about you. Knowing your home address does not mean the assistant is allowed to disclose it.
- Friday: action. You previously asked for a reminder before a delivery. The assistant must recognize the relevant moment and act, not merely answer a later question about the reminder.
- Later: deletion and recovery. You request deletion, then the service restores a backup. Information promised to be erased must not silently reappear.
A correct answer to “What is my current address?” tests only part of this story. Permission checks, timely action and deletion after recovery need their own observable outcomes. Testing each operation separately also misses interactions: a restore can undo a deletion, and a correction can leave stale information in a derived index.
| Test layer | What success can show | What it does not establish alone |
|---|---|---|
| Retrieval | The relevant source was found within the retrieval limit. | The answer is correct or disclosure is authorized. |
| Answer quality | The generated answer satisfies the task's scoring rule. | The underlying store is durable or an erased record is physically absent. |
| Agent behavior | The agent completes a task or acts at the required time. | Which part of the outcome came from memory rather than planning, tools or the language model. |
| Operational guarantees | Specified isolation, recovery, deletion or interface checks pass. | The system understands and uses memory well in open-ended tasks. |
What existing benchmarks do well—and where interpretation stops
It would be inaccurate to describe the entire field as simple fact recall. Existing work already covers reasoning, updates, learning, agent behavior and controlled access. The limitations below describe what a result should not be stretched to mean; they are not claims that these projects failed their own goals.
LongMemEval: remembering across conversations
LongMemEval tests information extraction, reasoning across sessions, knowledge updates, temporal reasoning and abstention. It supplies timestamped histories and supports evaluating outputs from your own system. A custom memory architecture is therefore not inherently incompatible. Those conversational results do not by themselves certify backup recovery, physical erasure or every application interface. Official code, data instructions and evaluation protocol.
MemoryAgentBench: incremental memory and learning
MemoryAgentBench organizes evaluation around accurate retrieval, test-time learning, long-range understanding and conflict resolution. This is broader than looking up facts. Its task-specific metrics should remain visible rather than being treated as a universal guarantee of operational reliability. Official implementation and task definitions.
MemoryArena: memory used in agent activity
MemoryArena evaluates memory in interdependent, multi-session agent tasks. That addresses an important gap between answering questions and using past experience to do work. Interpreting such outcomes still requires disclosing the agent, tools and model: their behavior contributes to the result alongside memory. Official project repository.
GateMem: access and changing memory
GateMem explicitly studies controlled memory access, updates and active forgetting across multiple principals. It is a counterexample to the claim that existing benchmarks ignore permissions or forgetting. Our further requirement is to distinguish observable answer-level behavior from physical erasure across storage, indexes and backups; one must not be inferred from the other. Official project website.
These are examples, not an exhaustive survey. The benchmark catalog covers additional families. The benchmark crosswalk connects their specific testing elements to our capability and module definitions. Missing mappings and incomplete integrations remain visible. An adapter we have not finished is a Mnemetric implementation gap, not a flaw in the upstream benchmark.
Why results can be hard to compare
Two percentages can look comparable while describing different experiments. Dataset revisions, session boundaries, timestamps, retrieval limits, language models, prompts, judges, allowed tools and resource budgets all affect the task. Supplying answer-bearing metadata can also make a task easier without making the memory better.
Our adapters must translate interfaces without silently changing the question being tested. Removing chronology from a temporal task is a broken integration. Giving one system extra hints is an unfair comparison. A system-specific task outside the original protocol can be useful, but it must be published separately with its own name and rules.
Averages also hide tradeoffs. Strong recall may coexist with poor abstention, unauthorized disclosure or excessive cost. Mnemetric's intended comparison is a profile of measured capabilities and explicit missing evidence, not a single score that lets a strength cancel out a critical failure.
Where Mnemosyne fits—and how it differs from Mnemetric
Mnemosyne is the memory system. Its architecture retains content-addressed evidence and builds searchable projections from it, with hybrid retrieval, time-aware beliefs, provenance and SQLite/PostgreSQL deployment paths. The memory systems page links the source snapshot and places this design beside other architectures.
Mnemetric is the evaluation project and website. It should measure Mnemosyne and other memory systems under disclosed, appropriate protocols. Its scope includes capabilities motivated by Mnemosyne's design, but its acceptance rules must describe observable behavior rather than require competitors to copy that design.
For Mnemosyne, finding a source is only one relevant check. We also need evidence that a correction affects the right state, provenance remains usable, tenant boundaries hold and supported interfaces behave consistently. These are reasons to test more, not reasons to discount a poor result on a valid existing task.
Implementation, development checks and benchmark results are different evidence levels. Some lifecycle requirements remain unfinished or only narrowly checked, including full erasure, broader recovery scenarios, complete working-memory behavior and multimodal coverage. The coverage status and development evidence identify the current boundaries. This page does not claim that Mnemosyne uniquely possesses these features or outperforms competing systems.
What Mnemetric must demonstrate
- Preserve official benchmarks. Use their official code, datasets, scoring and prescribed setups. Pin versions and disclose deviations. Executing an official scoring function on retained outputs is narrower evidence than completing the official end-to-end protocol.
- Map coverage explicitly. Connect each task, adapter and development check to the behavior it actually exercises. The proposed scope currently names 24 capabilities and 20 modules; these are coverage targets, not 20 completed benchmark suites.
- Test the full lifecycle and its interactions. Add inspectable scenarios for capture, retrieval, correction, learning, action, access, deletion and recovery, with clear pass conditions and resource measurements.
- Make results inspectable. Retain configurations, traces, failures and provenance. Show unmeasured capabilities as unmeasured. Separate synthetic development checks, adapted experiments and official benchmark runs.
- Earn comparisons. Run competing systems under the same applicable rules and disclose unsupported features and setup differences. Independent reproduction is needed to strengthen operator-run evidence.
The site is operated by the project behind Mnemosyne, which creates a conflict of interest that must stay visible. A broader test plan is not proof of a better memory system. The value of Mnemetric will depend on completed tests, reproducible results and honest treatment of failures—including Mnemosyne's.
Read the measurement and publication rules, inspect available results, or follow a requirement through the development-check mapping.