System landscape

Memory systems, side by side.

Does it find the right fact, how fast, and at what cost? Published numbers on one screen, each labelled with who measured it.

Measured by this harnessThird-party studyVendor self-reportNo-memory baselineNo honest number yet
Mnemosyne Recall@528.1%LongMemEval, 500 questions, signed bundle
Systems with a published figure21See the claims page for each source
Zep, measured by itself+9.2 ptsversus the same system run by Mem0
Head-to-head runs here0The open work for this site

Accuracy and speed

One study ran every system here, so these bars are comparable with each other. Mem0 wrote the study.

Answer accuracy

LoCoMo · higher is better

Run by Mem0
0%25%50%75%100%72.9%Full context (no memory)68.44%Mem0 graph66.88%Mem065.99%Zep58.1%LangMem52.9%OpenAI memory48.38%A-Memn/aMnemosyne
LoCoMo LLM-judged accuracy from one study authored by Mem0, higher is better
Full context (no memory)72.9%
Mem0 graph68.44%
Mem066.88%
Zep65.99%
LangMem58.1%
OpenAI memory52.9%
A-Mem48.38%
Mnemosyneno measurement

From the Mem0 paper, which Mem0 wrote and which scores its rivals. Zep disputes its number. We skip LoCoMo: its answer key is unreliable.

Speed

Seconds per query · lower is better

Run by Mem0
0s5s10s15s20s0.47 sOpenAI memory0.71 sMem01.09 sMem0 graph1.29 sZep1.41 sA-Mem9.87 sFull context (no memory)18.53 sLangMemn/aMnemosyne
Typical total latency per query, lower is better
OpenAI memory0.47 s
Mem00.71 s
Mem0 graph1.09 s
Zep1.29 s
A-Mem1.41 s
Full context (no memory)9.87 s
LangMem18.53 s
Mnemosyneno measurement

A typical query, search and answer together. We have not timed Mnemosyne yet.

Slow tail

Worst case, seconds · lower is better

Run by Mem0
0s5s10s15s20s0.89 sOpenAI memory1.44 sMem02.59 sMem0 graph2.93 sZep4.37 sA-Mem17.12 sFull context (no memory)60.40 s ↑LangMemn/aMnemosyne
Slow tail total latency per query, lower is better
OpenAI memory0.89 s
Mem01.44 s
Mem0 graph2.59 s
Zep2.93 s
A-Mem4.37 s
Full context (no memory)17.12 s
LangMem60.40 s
Mnemosyneno measurement

The slowest 5% of queries. LangMem hits 60 s, past the edge of the chart.

What projects claim about themselves

Each ran its own setup, so these compare with nothing else.

Every claim

Self-reported LoCoMo

LoCoMo · higher is better

Self-reported
0%25%50%75%100%96.1%ByteRover92.5%Mem092%Hindsight88.83%MemOS88.24%CORE87%Memori83.05%Nemori75.14%Zep
Self-reported LoCoMo accuracy from each project
ByteRover96.1%
Mem092.5%
Hindsight92%
MemOS88.83%
CORE88.24%
Memori87%
Nemori83.05%
Zep75.14%

From each project’s own README or blog. Different judges, different setups.

Same system, two operators

Zep on LoCoMo · higher is better

Disputed
0%25%50%75%100%65.99%Run by Mem075.14%Run by Zep
Zep LoCoMo accuracy measured by Mem0 and by Zep
Run by Mem065.99%
Run by Zep75.14%

Same software, two scorekeepers, nine points apart. This is exactly why one neutral operator should run everything.

LongMemEval accuracy, Zep paper

LLM-judged · higher is better

Self-reported
0%25%50%75%100%71.2%Zep gpt-4o60.2%Full context gpt-4o63.8%Zep gpt-4o-mini55.4%Full context gpt-4o-mini
LongMemEval accuracy from the Zep paper, two readers
Zep gpt-4o71.2%
Full context gpt-4o60.2%
Zep gpt-4o-mini63.8%
Full context gpt-4o-mini55.4%

Each reader has its own no-memory baseline. Compare Zep only with the baseline beside it.

Measured by us

Only one question here: was the right evidence in the top five? Nothing is graded by a language model, so these are not accuracy scores.

LongMemEval retrieval

Recall@5 and nDCG@5 · higher is better

This harness
0%25%50%75%100%29.67%nDCG@528.06%Recall@5
Mnemosyne LongMemEval retrieval with reported intervals
nDCG@529.67% (interval 26.38% to 33.04%)
Recall@528.06% (interval 24.88% to 31.37%)

500 questions, signed and reproducible. Other projects report much higher recall, but count hits differently — see the caveats. Inspect the run

Multi-hop retrieval

HippoRAG suite, Recall@5 · higher is better

This harness
0%25%50%75%100%37.4%Mnemosyne HotpotQA96.3%HippoRAG 2 HotpotQA23.73%Mnemosyne 2WikiMultiHopQA90.4%HippoRAG 2 2WikiMultiHopQA10.42%Mnemosyne MuSiQue74.7%HippoRAG 2 MuSiQue
Mnemosyne versus published HippoRAG 2 Recall@5 on three multi-hop datasets
Mnemosyne HotpotQA37.4%
HippoRAG 2 HotpotQA96.3%
Mnemosyne 2WikiMultiHopQA23.73%
HippoRAG 2 2WikiMultiHopQA90.4%
Mnemosyne MuSiQue10.42%
HippoRAG 2 MuSiQue74.7%

1,000 questions each. Red is HippoRAG 2’s published result: the bar to beat, not a run here. Our graph search did not fire on a single query, which is why these are low.

Capabilities at a glance

What each system can do, not how well it scored.

SystemOpen sourceSelf-hostedGraph memoryTime-aware factsProvenanceMCP serverBest known for
Mnemosyne Operator entryYesYesYesPartialYesYesCalibrated abstention, prospective actions, auditable evidence
Mem0YesYesPartialLimitedLimitedYesSimple API, wide adoption
Zep / GraphitiPartialPartialYesYesPartialYesTemporal knowledge graph
LettaYesYesNoLimitedLimitedYesAgent-managed memory blocks
LangMemYesYesNoLimitedLimitedNoNative fit for LangGraph agents
CogneeYesYesYesPartialPartialYesDocuments to knowledge graph pipelines
MemOSYesYesYesPartialPartialPartialMemory as an operating-system layer
supermemoryPartialPartialPartialLimitedLimitedYesDrop-in memory API and connectors
OpenAI memoryNoNoNoNoNoNoBuilt into ChatGPT, no setup, no control
YesPartialLimited: exists, not a design centreNo

Our own summary, deliberately rough and not verified. Spotted something wrong? Tell us.

What turns this page into a benchmark

Until these four steps land, grey and striped bars are citations. Only violet bars are measurements.

  1. Run every system hereMem0, Zep, Letta, LangMem, Cognee and MemOS through the same LongMemEval harness that produced the violet bars.
  2. Add judged accuracyA separate column with the reader and judge model named, never mixed with recall.
  3. Time it on one machinep50 and p95 for every system, next to cost per query.
  4. Sign and ship the bundleEvery run in the append-only ledger, with a one-command reproduction.