All research

Benchmark report 01

August 26, 20268 min read

Measuring long-term conversational memory on LoCoMo

A transparent end-to-end evaluation of Copepod's ability to retrieve and answer questions from long, multi-session conversations.

Headline result

92.34%

latest core-run answer accuracy

1,422 of 1,540 answers were judged correct in the latest published core run.

What this result means

This is an end-to-end result: Copepod selected evidence, an answer model produced the response, and a separate judge evaluated it. It measures the complete system path, not retrieval alone.

Our evaluation philosophy

Run our own race. Measure it honestly.

We evaluate Copepod as a complete system: a modern answer model working with a durable memory layer. We are not trying to reproduce another memory company's architecture or tune our work to match a competitor's reported score. We use capable modern models deliberately, build the memory system we believe agents need, and measure whether it improves over time.

Quality and efficiency move together. We report answer and judge tokens alongside accuracy because a memory system should improve the useful context available to an agent without treating an ever-larger prompt as the answer.

Every published score keeps its protocol, category breakdown, and policy boundary attached. We show genuine core runs; benchmark-only fallbacks and retrieval-only diagnostics do not become headline performance claims.

Published core-run history

Progress by run, with the categories intact

Benchmark-only fallback variants are excluded. Every point below is a full 1,540-question end-to-end core run.

Overall answer accuracy

Same LoCoMo denominator; interpret protocol changes beside the line.

+2.01 pp
88%90%92%94%90.32%Run 191.43%Run 292.34%Run 3
RunOverallMulti-hopTemporalOpen-domainSingle-hop

Extractive v19 · adaptive high recall

Historical baseline

90.32%

1391 / 1540

89.72%

253 / 282

88.47%

284 / 321

68.75%

66 / 96

93.70%

788 / 841

Candidate auto · adaptive midpoint

August 25, 2026

91.43%

1408 / 1540

92.20%

260 / 282

91.59%

294 / 321

65.63%

63 / 96

94.05%

791 / 841

Temporal-win checkpoint · adaptive midpoint

August 26, 2026

92.34%

1422 / 1540

92.91%

262 / 282

93.15%

299 / 321

70.83%

68 / 96

94.29%

793 / 841

Token usage

Cost per evaluated question

Totals include full answer generation and judging across all 1,540 questions; category totals are not separately tokenised.

Extractive v19 · adaptive high recall

14.78M total

9,598tokens / question

Answering
11.43M
Judging
3.35M

Candidate auto · adaptive midpoint

13.19M total

8,564tokens / question

Answering
9.84M
Judging
3.35M

Temporal-win checkpoint · adaptive midpoint

13.73M total

8,918tokens / question

Answering
10.38M
Judging
3.35M

Extractive v19 · adaptive high recall

Historical core baseline. Exact frozen artifact paths are unavailable in the current checkout.

Context policy: extractive-v19 adaptive-high-recall

Candidate auto · adaptive midpoint

Completed product-visible core run. Its lower open-domain result failed a later explicit gate, so it is reported as an observed core result rather than acceptance evidence.

Context policy: candidate-auto-adaptive-midpoint

Temporal-win checkpoint · adaptive midpoint

User-accepted temporal-win checkpoint from a fresh, commit-bound full run. Every LoCoMo category improved versus the historical baseline; the observed delta also includes compiler, context-policy, and stochastic model-call changes.

Context policy: candidate-2b38545 auto-adaptive-midpoint; bounded product-visible supporting sessions

Scorecard

What we measured—and how to read it

92.34% · 1,422 / 1,540

Answer accuracy

The proportion of benchmark questions whose final answer was judged correct.

What it is for

Use this to compare complete question-answering performance when the benchmark version, answer model, judge, and prompt protocol match.

It cannot identify whether an error came from retrieval, evidence assembly, the answer model, or the judge.

4 LoCoMo categories

Category breakdown

Multi-hop, temporal, open-domain, and single-hop scores are reported separately for every published core run.

What it is for

Use it to see where an improvement occurred instead of treating the overall result as a single undifferentiated number.

Category deltas are diagnostic; changes to the context policy mean they are not isolated causal claims.

Reported separately

Retrieval quality

Whether supporting evidence was found within the configured retrieval depth before answer generation.

What it is for

Use it to diagnose memory search and ranking independently of answer generation.

A retrieval hit does not prove the final answer is correct, especially when a temporal answer loses its resolved date.

Run-specific

Latency

Time spent executing the evaluated path, reported with runtime and isolation conditions.

What it is for

Use it for capacity and experience planning only when hardware, tenancy, load, and timing boundaries are comparable.

Latency from a recovered or shared runtime is not an authoritative performance claim.

Run protocol

Benchmark
LoCoMo conversational QA
Question denominator
1,540
System under test
Copepod retrieval + answer generation
Answer model
GPT-5.6 Luna
Primary judge
GPT-5.6 Sol
Measurement type
End-to-end QA accuracy

What this supports

  • How reliably the evaluated system answers this benchmark's questions across long conversations.
  • How the end-to-end result changes when the evaluator changes.
  • The exact denominator, models, and evaluation boundary behind the claim.

What it does not support

  • A universal ranking across other datasets, versions, or vendor protocols.
  • Retrieval-only quality, oracle-context quality, or an isolated model capability unless separately reported.
  • Production latency under every tenant, model, hardware, or concurrent workload.