System landscape
Memory systems, side by side.
Does it find the right fact, how fast, and at what cost? Published numbers on one screen, each labelled with who measured it.
Accuracy and speed
One study ran every system here, so these bars are comparable with each other. Mem0 wrote the study.
Answer accuracy
LoCoMo · higher is better
| Full context (no memory) | 72.9% |
|---|---|
| Mem0 graph | 68.44% |
| Mem0 | 66.88% |
| Zep | 65.99% |
| LangMem | 58.1% |
| OpenAI memory | 52.9% |
| A-Mem | 48.38% |
| Mnemosyne | no measurement |
From the Mem0 paper, which Mem0 wrote and which scores its rivals. Zep disputes its number. We skip LoCoMo: its answer key is unreliable.
Speed
Seconds per query · lower is better
| OpenAI memory | 0.47 s |
|---|---|
| Mem0 | 0.71 s |
| Mem0 graph | 1.09 s |
| Zep | 1.29 s |
| A-Mem | 1.41 s |
| Full context (no memory) | 9.87 s |
| LangMem | 18.53 s |
| Mnemosyne | no measurement |
A typical query, search and answer together. We have not timed Mnemosyne yet.
Slow tail
Worst case, seconds · lower is better
| OpenAI memory | 0.89 s |
|---|---|
| Mem0 | 1.44 s |
| Mem0 graph | 2.59 s |
| Zep | 2.93 s |
| A-Mem | 4.37 s |
| Full context (no memory) | 17.12 s |
| LangMem | 60.40 s |
| Mnemosyne | no measurement |
The slowest 5% of queries. LangMem hits 60 s, past the edge of the chart.
What projects claim about themselves
Each ran its own setup, so these compare with nothing else.
Self-reported LoCoMo
LoCoMo · higher is better
| ByteRover | 96.1% |
|---|---|
| Mem0 | 92.5% |
| Hindsight | 92% |
| MemOS | 88.83% |
| CORE | 88.24% |
| Memori | 87% |
| Nemori | 83.05% |
| Zep | 75.14% |
From each project’s own README or blog. Different judges, different setups.
Same system, two operators
Zep on LoCoMo · higher is better
| Run by Mem0 | 65.99% |
|---|---|
| Run by Zep | 75.14% |
Same software, two scorekeepers, nine points apart. This is exactly why one neutral operator should run everything.
LongMemEval accuracy, Zep paper
LLM-judged · higher is better
| Zep gpt-4o | 71.2% |
|---|---|
| Full context gpt-4o | 60.2% |
| Zep gpt-4o-mini | 63.8% |
| Full context gpt-4o-mini | 55.4% |
Each reader has its own no-memory baseline. Compare Zep only with the baseline beside it.
Measured by us
Only one question here: was the right evidence in the top five? Nothing is graded by a language model, so these are not accuracy scores.
LongMemEval retrieval
Recall@5 and nDCG@5 · higher is better
| nDCG@5 | 29.67% (interval 26.38% to 33.04%) |
|---|---|
| Recall@5 | 28.06% (interval 24.88% to 31.37%) |
500 questions, signed and reproducible. Other projects report much higher recall, but count hits differently — see the caveats. Inspect the run
Multi-hop retrieval
HippoRAG suite, Recall@5 · higher is better
| Mnemosyne HotpotQA | 37.4% |
|---|---|
| HippoRAG 2 HotpotQA | 96.3% |
| Mnemosyne 2WikiMultiHopQA | 23.73% |
| HippoRAG 2 2WikiMultiHopQA | 90.4% |
| Mnemosyne MuSiQue | 10.42% |
| HippoRAG 2 MuSiQue | 74.7% |
1,000 questions each. Red is HippoRAG 2’s published result: the bar to beat, not a run here. Our graph search did not fire on a single query, which is why these are low.
Capabilities at a glance
What each system can do, not how well it scored.
| System | Open source | Self-hosted | Graph memory | Time-aware facts | Provenance | MCP server | Best known for |
|---|---|---|---|---|---|---|---|
| Mnemosyne Operator entry | Yes | Yes | Yes | Partial | Yes | Yes | Calibrated abstention, prospective actions, auditable evidence |
| Mem0 | Yes | Yes | Partial | Limited | Limited | Yes | Simple API, wide adoption |
| Zep / Graphiti | Partial | Partial | Yes | Yes | Partial | Yes | Temporal knowledge graph |
| Letta | Yes | Yes | No | Limited | Limited | Yes | Agent-managed memory blocks |
| LangMem | Yes | Yes | No | Limited | Limited | No | Native fit for LangGraph agents |
| Cognee | Yes | Yes | Yes | Partial | Partial | Yes | Documents to knowledge graph pipelines |
| MemOS | Yes | Yes | Yes | Partial | Partial | Partial | Memory as an operating-system layer |
| supermemory | Partial | Partial | Partial | Limited | Limited | Yes | Drop-in memory API and connectors |
| OpenAI memory | No | No | No | No | No | No | Built into ChatGPT, no setup, no control |
Our own summary, deliberately rough and not verified. Spotted something wrong? Tell us.
What turns this page into a benchmark
Until these four steps land, grey and striped bars are citations. Only violet bars are measurements.
- Run every system hereMem0, Zep, Letta, LangMem, Cognee and MemOS through the same LongMemEval harness that produced the violet bars.
- Add judged accuracyA separate column with the reader and judge model named, never mixed with recall.
- Time it on one machinep50 and p95 for every system, next to cost per query.
- Sign and ship the bundleEvery run in the append-only ledger, with a one-command reproduction.