Extractive v19 · adaptive high recall
14.78M total
9,598tokens / question
- Answering
- 11.43M
- Judging
- 3.35M
Benchmark report 01
August 26, 20268 min readA transparent end-to-end evaluation of Copepod's ability to retrieve and answer questions from long, multi-session conversations.
Headline result
92.34%
latest core-run answer accuracy
1,422 of 1,540 answers were judged correct in the latest published core run.
What this result means
This is an end-to-end result: Copepod selected evidence, an answer model produced the response, and a separate judge evaluated it. It measures the complete system path, not retrieval alone.
Our evaluation philosophy
We evaluate Copepod as a complete system: a modern answer model working with a durable memory layer. We are not trying to reproduce another memory company's architecture or tune our work to match a competitor's reported score. We use capable modern models deliberately, build the memory system we believe agents need, and measure whether it improves over time.
Quality and efficiency move together. We report answer and judge tokens alongside accuracy because a memory system should improve the useful context available to an agent without treating an ever-larger prompt as the answer.
Every published score keeps its protocol, category breakdown, and policy boundary attached. We show genuine core runs; benchmark-only fallbacks and retrieval-only diagnostics do not become headline performance claims.
Published core-run history
Benchmark-only fallback variants are excluded. Every point below is a full 1,540-question end-to-end core run.
Overall answer accuracy
Same LoCoMo denominator; interpret protocol changes beside the line.
| Run | Overall | Multi-hop | Temporal | Open-domain | Single-hop |
|---|---|---|---|---|---|
Extractive v19 · adaptive high recall Historical baseline | 90.32% 1391 / 1540 | 89.72% 253 / 282 | 88.47% 284 / 321 | 68.75% 66 / 96 | 93.70% 788 / 841 |
Candidate auto · adaptive midpoint August 25, 2026 | 91.43% 1408 / 1540 | 92.20% 260 / 282 | 91.59% 294 / 321 | 65.63% 63 / 96 | 94.05% 791 / 841 |
Temporal-win checkpoint · adaptive midpoint August 26, 2026 | 92.34% 1422 / 1540 | 92.91% 262 / 282 | 93.15% 299 / 321 | 70.83% 68 / 96 | 94.29% 793 / 841 |
Token usage
Totals include full answer generation and judging across all 1,540 questions; category totals are not separately tokenised.
Extractive v19 · adaptive high recall
14.78M total
9,598tokens / question
Candidate auto · adaptive midpoint
13.19M total
8,564tokens / question
Temporal-win checkpoint · adaptive midpoint
13.73M total
8,918tokens / question
Extractive v19 · adaptive high recall
Historical core baseline. Exact frozen artifact paths are unavailable in the current checkout.
Context policy: extractive-v19 adaptive-high-recall
Candidate auto · adaptive midpoint
Completed product-visible core run. Its lower open-domain result failed a later explicit gate, so it is reported as an observed core result rather than acceptance evidence.
Context policy: candidate-auto-adaptive-midpoint
Temporal-win checkpoint · adaptive midpoint
User-accepted temporal-win checkpoint from a fresh, commit-bound full run. Every LoCoMo category improved versus the historical baseline; the observed delta also includes compiler, context-policy, and stochastic model-call changes.
Context policy: candidate-2b38545 auto-adaptive-midpoint; bounded product-visible supporting sessions
Scorecard
92.34% · 1,422 / 1,540
The proportion of benchmark questions whose final answer was judged correct.
What it is for
Use this to compare complete question-answering performance when the benchmark version, answer model, judge, and prompt protocol match.
It cannot identify whether an error came from retrieval, evidence assembly, the answer model, or the judge.
4 LoCoMo categories
Multi-hop, temporal, open-domain, and single-hop scores are reported separately for every published core run.
What it is for
Use it to see where an improvement occurred instead of treating the overall result as a single undifferentiated number.
Category deltas are diagnostic; changes to the context policy mean they are not isolated causal claims.
Reported separately
Whether supporting evidence was found within the configured retrieval depth before answer generation.
What it is for
Use it to diagnose memory search and ranking independently of answer generation.
A retrieval hit does not prove the final answer is correct, especially when a temporal answer loses its resolved date.
Run-specific
Time spent executing the evaluated path, reported with runtime and isolation conditions.
What it is for
Use it for capacity and experience planning only when hardware, tenancy, load, and timing boundaries are comparable.
Latency from a recovered or shared runtime is not an authoritative performance claim.