Published claims

What every memory system claims.

The scores these projects publish about themselves, collected with a link to where each one came from.

This is not a ranking
These projects each tested themselves, using different models, judges and settings. So this is a list, not a ranking, and small gaps mean nothing. Only violet rows were measured here.
Figures collected93With source, commit and line
Systems and projects22Open source and closed
Self-reported54Not independently reproduced
Measured here5By the signed harness
Measured by this harnessSelf-reported by the projectRun by another partyBaseline or no memory

These figures as charts

Every chart on this data lives on the front page, grouped by what it measures.

Open the charts

The projects

Licences come from GitHub. Closed products have no repository.

Mem0

Open-source SDK plus managed platform

Apache-2.013 claims

Platform scores include proprietary optimizations

Repository

MemPalace

Open source, local-first

MIT8 claims

Declines to compare itself with other projects

Repository

Hindsight

Open source plus cloud

MIT9 claims

Maker of the Agent Memory Benchmark

Repository

OpenViking

Open-source context database

AGPL-3.01 claim

Reports a range, no single figure

Repository

gbrain

Open source

MIT4 claims

Evaluated in a separate public eval repository

Repository

supermemory

Open source plus hosted

MIT4 claims

Claims #1 without a figure for answer quality

Repository

agentmemory

Open source

Apache-2.02 claims

Says only its own retrieval figure is measured

Repository

Memori

Memory infrastructure

No standard licence detected1 claim

LoCoMo paper arXiv 2603.19935

Repository

MemOS

Open-source memory OS

Apache-2.05 claims

Scores come from its own OmniMemEval

Repository

ByteRover

Memory layer for coding agents

No standard licence detected2 claims

LoCoMo run on its production codebase

Repository

CORE

Personal memory layer

No standard licence detected1 claim

Benchmark repo published separately

Repository

Memanto

Open source

MIT2 claims

Warns scores are not comparable across projects

Repository

Cognee

Open-source knowledge graph memory

Apache-2.04 claims

BEAM runs scoped to a few questions

Repository

ReMe

Open-source memory kit

Apache-2.03 claims
Repository

Nemori

Open source

MIT1 claim
Repository

Zep

Graphiti open source; Zep Cloud managed

Apache-2.0 (Graphiti)6 claims

Disputes the figure in the Mem0 paper

Repository

LightMem

Research memory framework

MIT0 claims

Its README compares several systems on one harness

Repository

Letta

Open source

Apache-2.00 claims

No benchmark figures in its README

Repository

Honcho

Memory library for stateful agents

AGPL-3.00 claims

Points to an evals page; no figures in its README

Repository

OpenAI memory

Closed, built into ChatGPT

Closed source3 claims

Only appears via the Mem0 paper

Mnemosyne

Operator entry of this site

Not recorded here5 claims

Held to the same rules as every system

Repository

Claims with no usable figure:

supermemory · claims #1 on LongMemEval, LoCoMo, ConvoMemLetta · no figures in READMEHoncho · evals page, no figures in READMEZep / Graphiti · no figures in READMEOpenViking · range only, 80 to 83% on LoCoMo

Every claim

Type to filter. Each row links to where the number came from.

Download JSON

SystemBenchmarkMetricValueFamilyWho ran itSource
HindsightBEAM 100KaccuracyAgent Memory Benchmark comparison table.86.2%Answer qualitySelf-reportedAgent Memory Benchmark, run by Vectorize (maker of Hindsight)
HindsightBEAM 1MaccuracyAgent Memory Benchmark comparison table.79.1%Answer qualitySelf-reportedAgent Memory Benchmark, run by Vectorize (maker of Hindsight)
CogneeBEAM 100Kscore (0 to 1 scale)Reported 0.79: four rounds over 20 questions from one held-out conversation.79%Answer qualitySelf-reportedCognee READMEcommit b32d8af line 291
Hindsight · single queryBEAM 100KaccuracyHindsight results page, single-query mode.75%Answer qualitySelf-reportedHindsight benchmarks page
Hindsight · single queryBEAM 1MaccuracyHindsight results page, single-query mode.73.9%Answer qualitySelf-reportedHindsight benchmarks page
CogneeBEAM 10Mscore (0 to 1 scale)Reported 0.67: exploratory, question-type routing selected and scored on the same questions.67%Answer qualitySelf-reportedCognee READMEcommit b32d8af line 292
ReMeBEAM 100Kagentic score20 cases, 400 questions.66.1%Answer qualitySelf-reportedReMe READMEcommit 084c02e line 356
ReMeBEAM 1Magentic score35 cases, 700 questions.65%Answer qualitySelf-reportedReMe READMEcommit 084c02e line 357
HindsightBEAM 10MaccuracyAgent Memory Benchmark and the Hindsight results page agree.64.1%Answer qualitySelf-reportedAgent Memory Benchmark, run by Vectorize (maker of Hindsight)
Mem0BEAM 1MscoreSame setup.64.1%Answer qualitySelf-reportedMem0 READMEcommit c93420c line 51
MemOSBEAM 10MscoreEvaluated through OmniMemEval.56.75%Answer qualitySelf-reportedMemOS READMEcommit a7367d0 line 82
Mem0BEAM 10MscoreSame setup.48.6%Answer qualitySelf-reportedMem0 READMEcommit c93420c line 52
MemPalaceConvoMemaverage recallAll categories, 250 items, 50 per category.92.9%Retrieval recallSelf-reportedMemPalace READMEcommit 35dc621 line 268
supermemoryConvoMemranking claim: #1States #1 with no figure in the README.—Metric not statedSelf-reportedsupermemory READMEcommit 3535ff7 line 355
MemOSHaluMemscoreEvaluated through OmniMemEval.80.91%Answer qualitySelf-reportedMemOS READMEcommit a7367d0 line 81
HippoRAG 2HippoRAG: 2WikiMultiHopQARecall@5Published upstream result, Llama-3.3-70B.90.4%Retrieval recallRun by another partyMnemetric Phase 11 evidence notecommit main line 35
MnemosyneHippoRAG: 2WikiMultiHopQARecall@51,000 questions, July 2026, retrieval only.23.73%Retrieval recallMeasured hereMnemetric 2WikiMultiHopQA retrieval reportcommit main
HippoRAG 2HippoRAG: HotpotQARecall@5Published upstream result, Llama-3.3-70B.96.3%Retrieval recallRun by another partyMnemetric Phase 11 evidence notecommit main line 36
MnemosyneHippoRAG: HotpotQARecall@51,000 questions, July 2026, retrieval only.37.4%Retrieval recallMeasured hereMnemetric HotpotQA retrieval reportcommit main
HippoRAG 2HippoRAG: MuSiQueRecall@5Published upstream result, Llama-3.3-70B.74.7%Retrieval recallRun by another partyMnemetric Phase 11 evidence notecommit main line 34
MnemosyneHippoRAG: MuSiQueRecall@51,000 questions, July 2026, retrieval only.10.42%Retrieval recallMeasured hereMnemetric MuSiQue retrieval reportcommit main
HindsightLifeBench enaccuracyAgent Memory Benchmark.71.5%Answer qualitySelf-reportedAgent Memory Benchmark, run by Vectorize (maker of Hindsight)
hybrid-searchLifeBench enaccuracyAgent Memory Benchmark baseline: plain hybrid search.61%Answer qualityRun by another partyAgent Memory Benchmark, run by Vectorize (maker of Hindsight)
ByteRoverLoCoMooverall accuracy1,982 questions, 272 documents; production codebase, no separate prototype.96.1%Answer qualitySelf-reportedByteRover READMEcommit 1052ac1 line 55
Mem0LoCoMoscoreApril 2026 algorithm on the managed platform, which includes proprietary optimizations; single pass, top_200; was 71.4.92.5%Answer qualitySelf-reportedMem0 READMEcommit c93420c line 49
HindsightLoCoMoaccuracyAgent Memory Benchmark; 1,986 questions.92%Answer qualitySelf-reportedAgent Memory Benchmark, run by Vectorize (maker of Hindsight)
MemOSLoCoMoscoreEvaluated through OmniMemEval, a framework run by the same organisation.88.83%Answer qualitySelf-reportedMemOS READMEcommit a7367d0 line 78
CORELoCoMoaverage accuracyAcross single-hop, multi-hop, open-domain and temporal questions.88.24%Answer qualitySelf-reportedCORE READMEcommit 4a5b18d line 244
MemoriLoCoMooverall accuracyAbout 721 tokens per query, 2.8% of the full-context footprint.87%Answer qualitySelf-reportedMemori READMEcommit 574b1ea line 130
NemoriLoCoMoLLM scoreReported 0.8305 overall, version V5; 1,540 questions across the four categories.83.05%Answer qualitySelf-reportedNemori READMEcommit d2a6dff line 171
CogneeLoCoMoaccuracyAgent Memory Benchmark, run by Hindsight's maker.80.3%Answer qualityRun by another partyAgent Memory Benchmark, run by Vectorize (maker of Hindsight)
hybrid-searchLoCoMoaccuracyAgent Memory Benchmark baseline: plain hybrid search.79.1%Answer qualityRun by another partyAgent Memory Benchmark, run by Vectorize (maker of Hindsight)
ZepLoCoMoJ score (LLM judge)Zep corrected its own earlier post and reports 75.14 +/- 0.17, against 65.99 in the Mem0 paper.75.14%Answer qualitySelf-reportedZep blog, corrected LoCoMo result
FullTextLoCoMoACCLightMem authors; backbone and judge gpt-4o-mini.73.83%Answer qualityRun by another partyLightMem READMEcommit 8449d57 line 393
Full context (no memory)LoCoMoJ score (LLM judge)Whole conversation in the prompt, about 26,000 tokens.72.9%Answer qualityRun by another partyMem0 paper, arXiv 2504.19413
Mem0 graphLoCoMoJ score (LLM judge)Mem0 paper.68.44%Answer qualitySelf-reportedMem0 paper, arXiv 2504.19413
Mem0LoCoMoJ score (LLM judge)Mem0 paper.66.88%Answer qualitySelf-reportedMem0 paper, arXiv 2504.19413
ZepLoCoMoJ score (LLM judge)Mem0 authors running Zep.65.99%Answer qualityRun by another partyMem0 paper, arXiv 2504.19413
A-MemLoCoMoACCLightMem authors; backbone and judge gpt-4o-mini.64.16%Answer qualityRun by another partyLightMem READMEcommit 8449d57 line 395
NaiveRAGLoCoMoACCLightMem authors; backbone and judge gpt-4o-mini.63.64%Answer qualityRun by another partyLightMem READMEcommit 8449d57 line 394
Mem0 · APILoCoMoACCLightMem authors; the hosted API.61.69%Answer qualityRun by another partyLightMem READMEcommit 8449d57 line 399
Mem0 graph · APILoCoMoACCLightMem authors; the hosted API with graph.60.32%Answer qualityRun by another partyLightMem READMEcommit 8449d57 line 400
MemoryOS · eval buildLoCoMoACCLightMem authors; the evaluation build.58.25%Answer qualityRun by another partyLightMem READMEcommit 8449d57 line 396
LangMemLoCoMoJ score (LLM judge)Mem0 paper.58.1%Answer qualityRun by another partyMem0 paper, arXiv 2504.19413
MemoryOS · PyPILoCoMoACCLightMem authors; the PyPI release.54.87%Answer qualityRun by another partyLightMem READMEcommit 8449d57 line 397
OpenAI memoryLoCoMoJ score (LLM judge)OpenAI memory as run by the Mem0 authors; closed source.52.9%Answer qualityRun by another partyMem0 paper, arXiv 2504.19413
A-MemLoCoMoJ score (LLM judge)Mem0 paper.48.38%Answer qualityRun by another partyMem0 paper, arXiv 2504.19413
Mem0 · open sourceLoCoMoACCLightMem authors; the open-source package.36.49%Answer qualityRun by another partyLightMem READMEcommit 8449d57 line 398
OpenVikingLoCoMoaccuracyReports 80 to 83% across three agent integrations, against 24 to 57% on their native memory; reader Doubao 2.0 Pro.80 to 83%Answer qualitySelf-reportedOpenViking READMEcommit 10f3681 line 120
LangMemLoCoMototal latency p95Search alone was 59.82 s.60.4 sLatencyRun by another partyMem0 paper, arXiv 2504.19413
LangMemLoCoMototal latency p50Search alone was 17.99 s; the total is search plus answer.18.53 sLatencyRun by another partyMem0 paper, arXiv 2504.19413
Full context (no memory)LoCoMototal latency p95Whole conversation in the prompt.17.117 sLatencyRun by another partyMem0 paper, arXiv 2504.19413
Full context (no memory)LoCoMototal latency p50Whole conversation in the prompt.9.87 sLatencyRun by another partyMem0 paper, arXiv 2504.19413
A-MemLoCoMototal latency p95Search plus answer, seconds.4.374 sLatencyRun by another partyMem0 paper, arXiv 2504.19413
ZepLoCoMototal latency p95Search plus answer, seconds.2.926 sLatencyRun by another partyMem0 paper, arXiv 2504.19413
Mem0 graphLoCoMototal latency p95Search plus answer, seconds.2.59 sLatencySelf-reportedMem0 paper, arXiv 2504.19413
Mem0LoCoMototal latency p95Search plus answer, seconds.1.44 sLatencySelf-reportedMem0 paper, arXiv 2504.19413
A-MemLoCoMototal latency p50Search plus answer, seconds.1.41 sLatencyRun by another partyMem0 paper, arXiv 2504.19413
ZepLoCoMototal latency p50Search plus answer, seconds.1.292 sLatencyRun by another partyMem0 paper, arXiv 2504.19413
Mem0 graphLoCoMototal latency p50Search plus answer, seconds.1.091 sLatencySelf-reportedMem0 paper, arXiv 2504.19413
OpenAI memoryLoCoMototal latency p95No search step.0.889 sLatencyRun by another partyMem0 paper, arXiv 2504.19413
Mem0LoCoMototal latency p50Search plus answer, seconds.0.708 sLatencySelf-reportedMem0 paper, arXiv 2504.19413
OpenAI memoryLoCoMototal latency p50No search step; memories are extracted manually in the prompt.0.466 sLatencyRun by another partyMem0 paper, arXiv 2504.19413
MemPalace · hybrid v5LoCoMoR@10Hybrid v5, top-10, no rerank; same 1,986 questions.88.9%Retrieval recallSelf-reportedMemPalace READMEcommit 35dc621 line 267
MemPalace · sessionLoCoMoR@10Session level, top-10, no rerank; 1,986 questions.60.3%Retrieval recallSelf-reportedMemPalace READMEcommit 35dc621 line 266
MemantoLoCoMoscoreSame caveat.87.1%Metric not statedSelf-reportedMemanto READMEcommit aac858c line 332
supermemoryLoCoMoranking claim: #1States #1 with no figure in the README.—Metric not statedSelf-reportedsupermemory READMEcommit 3535ff7 line 354
HindsightLongMemEval SaccuracyAgent Memory Benchmark, operated by Vectorize, which makes Hindsight.94.6%Answer qualitySelf-reportedAgent Memory Benchmark, run by Vectorize (maker of Hindsight)
Mem0LongMemEvalscoreSame setup; was 67.8.94.4%Answer qualitySelf-reportedMem0 READMEcommit c93420c line 50
ByteRoverLongMemEval Soverall accuracy500 questions, 23,867 documents.92.8%Answer qualitySelf-reportedByteRover READMEcommit 1052ac1 line 57
gbrainLongMemEvalaccuracy453 of 500; house reader with reranker.90.6%Answer qualitySelf-reportedgbrain-evals READMEcommit 48dd47b line 50
ReMeLongMemEval cleaned-sagentic score500 questions.89.4%Answer qualitySelf-reportedReMe READMEcommit 084c02e line 355
gbrain · gpt-5.4 readerLongMemEvalaccuracy447 of 500; gpt-5.4 reader on gbrain retrieval.89.4%Answer qualitySelf-reportedgbrain-evals READMEcommit 48dd47b line 51
MemOSLongMemEvalscoreEvaluated through OmniMemEval.89.2%Answer qualitySelf-reportedMemOS READMEcommit a7367d0 line 79
hybrid-searchLongMemEval SaccuracyAgent Memory Benchmark baseline: plain hybrid search.74%Answer qualityRun by another partyAgent Memory Benchmark, run by Vectorize (maker of Hindsight)
Zep · gpt-4o readerLongMemEval SaccuracyZep paper; reader gpt-4o; full-context baseline 60.2; latency 2.58 s against 28.9 s.71.2%Answer qualitySelf-reportedZep paper, arXiv 2501.13956, Table 2
Zep · gpt-4o-mini readerLongMemEval SaccuracyZep paper; reader gpt-4o-mini; full-context baseline 55.4; latency 3.20 s against 31.3 s.63.8%Answer qualitySelf-reportedZep paper, arXiv 2501.13956, Table 2
Full context (no memory) · gpt-4o readerLongMemEval SaccuracyZep paper; reader gpt-4o; about 115k tokens of context.60.2%Answer qualitySelf-reportedZep paper, arXiv 2501.13956, Table 2
Full context (no memory) · gpt-4o-mini readerLongMemEval SaccuracyZep paper; reader gpt-4o-mini; about 115k tokens of context.55.4%Answer qualitySelf-reportedZep paper, arXiv 2501.13956, Table 2
MemPalace · hybrid v4, held outLongMemEvalR@5Hybrid v4, held-out 450 questions; tuned on 50 dev questions.98.4%Retrieval recallSelf-reportedMemPalace READMEcommit 35dc621 line 246
MemPalace · rawLongMemEvalR@5500 questions; raw semantic search, no LLM, no heuristics.96.6%Retrieval recallSelf-reportedMemPalace READMEcommit 35dc621 line 245
gbrainLongMemEvalstrict recall_all@5451 of 470 questions; every required session must be in the top five; opaque session ids; with Voyage reranker.95.96%Retrieval recallSelf-reportedgbrain-evals READMEcommit 48dd47b line 49
agentmemoryLongMemEval SR@5500 questions; all-MiniLM-L6-v2 embeddings, local.95.2%Retrieval recallSelf-reportedagentmemory READMEcommit 007a1a7 line 356
supermemoryLongMemEvalRecall@15Adds about 720 tokens of context, a 99.4% reduction.95%Retrieval recallSelf-reportedsupermemory READMEcommit 3535ff7 line 357
gbrain · no rerankerLongMemEvalstrict recall_all@5434 of 470; same without the reranker.92.34%Retrieval recallSelf-reportedgbrain-evals READMEcommit 48dd47b line 88
MemPalace · hybrid + LLM rerank, recountedLongMemEvalstrict recall_all@5gbrain authors recounting MemPalace's saved rankings strictly (423 of 470); hybrid search plus LLM rerank.90%Retrieval recallRun by another partygbrain-evals READMEcommit 48dd47b line 89
BM25-only fallbackLongMemEval SR@5agentmemory's own keyword-only baseline on the same questions.86.2%Retrieval recallSelf-reportedagentmemory READMEcommit 007a1a7 line 357
MemPalace · raw, recountedLongMemEvalstrict recall_all@5Recount of the raw vector search (403 of 470).85.7%Retrieval recallRun by another partygbrain-evals READMEcommit 48dd47b line 91
MnemosyneLongMemEvalnDCG@5500 questions, interval 26.38 to 33.04.29.67%Retrieval recallMeasured hereMnemetric LongMemEval retrieval reportcommit main
MnemosyneLongMemEvalRecall@5500 questions, interval 24.88 to 31.37.28.06%Retrieval recallMeasured hereMnemetric LongMemEval retrieval reportcommit main
MemantoLongMemEvalscoreREADME calls these public recall benchmarks but does not state the metric, and warns cross-project scores are not comparable.89.8%Metric not statedSelf-reportedMemanto READMEcommit aac858c line 332
supermemoryLongMemEvalranking claim: #1States #1 with no figure in the README.—Metric not statedSelf-reportedsupermemory READMEcommit 3535ff7 line 353
MemPalaceMemBenchR@5ACL 2025, 8,500 items, all categories.80.3%Retrieval recallSelf-reportedMemPalace READMEcommit 35dc621 line 269
HindsightPersonaMem 32kaccuracyAgent Memory Benchmark.86.6%Answer qualitySelf-reportedAgent Memory Benchmark, run by Vectorize (maker of Hindsight)
hybrid-searchPersonaMem 32kaccuracyAgent Memory Benchmark baseline: plain hybrid search.84.4%Answer qualityRun by another partyAgent Memory Benchmark, run by Vectorize (maker of Hindsight)
CogneePersonaMem 32kaccuracyAgent Memory Benchmark, run by Hindsight's maker.81.8%Answer qualityRun by another partyAgent Memory Benchmark, run by Vectorize (maker of Hindsight)
MemOSPersonaMem v2scoreEvaluated through OmniMemEval.40.58%Answer qualitySelf-reportedMemOS READMEcommit a7367d0 line 80
Corrections we have made

Numbers that were wrong or unsourced on this site, and what replaced them.

  • Zep LongMemEval: an earlier version of the landscape page mixed the two readers in the Zep paper. The paper reports 63.8 with gpt-4o-mini (full context 55.4) and 71.2 with gpt-4o (full context 60.2).
  • Zep latency: an earlier version used the Zep paper's latency in a chart of the Mem0 paper. The Mem0 paper measures Zep at 1.292 s p50 and 2.926 s p95 in total.
  • LangMem latency: the figure shown was search time only (17.99 s and 59.82 s). The total is 18.53 s and 60.40 s.
  • Removed: self-reported LoCoMo figures for supermemory (81.6) and MemOS (73.3) that could not be traced to a source. MemOS now reports 88.83; supermemory publishes a #1 claim without a figure.
  • The capability table is an editorial summary and is labelled that way. Licences on this page come from GitHub's licence detection.

Collected 2026-10-06. Sources can change after that date. How the benchmarks differ · The charts