Docs · Getting started

Reading results

A score is useful only when you can see how it was produced. This page explains the charts, intervals and labels used across the site.

Recall and accuracy are different

Retrieval recall

Was the right memory among the first k returned? Deterministic: no language model is involved, so the same bundle always gives the same number.

Judged answer accuracy

Did a language model answer correctly from what was retrieved, as graded by another model? It depends on the reader and the judge, which must be named.

The two are never placed on one axis. A low recall and a high accuracy can both be true for the same system, because they measure different stages.

What the bar styles mean

VioletMeasured by this harness, signed and reproducible
Solid grey or colourOne third-party study; comparable within that chart
StripedA vendor’s own number; comparable with nothing else
Light greyNo-memory baseline, the whole history in the prompt
Dashed boxNo honest number exists yet
WhiskersThe interval reported by the run itself

Intervals and ties

Whiskers show the range a run reported for itself. If two systems’ ranges overlap, call it a tie.

Compare like with like

Two numbers only compare if the dataset, split, models and budgets match. The Compare runs page groups runs that match and refuses to rank across groups.

Inspect a run

Open any result to see its digests, publication status and per-question traces.