All research

Benchmark report 02

August 26, 20266 min read

Building confidence on LongMemEval before the full suite

A transparent development timeline on a frozen 12-question diagnostic set, used to improve temporal and knowledge-update memory before spending a full 500-question evaluation.

Headline result

75.00%

best fixed-12 diagnostic accuracy

9 of 12 answers were judged correct at the accepted diagnostic checkpoint.

What this result means

This is a development diagnostic, not an official LongMemEval score. We are deliberately not running or publishing the full 500-question suite until the retrieval and context compiler are stable enough for the result to be informative, reproducible, and worth the evaluation cost.

Our evaluation philosophy

Run our own race. Measure it honestly.

We evaluate Copepod as a complete system: a modern answer model working with a durable memory layer. We are not trying to reproduce another memory company's architecture or tune our work to match a competitor's reported score. We use capable modern models deliberately, build the memory system we believe agents need, and measure whether it improves over time.

Quality and efficiency move together. We report answer and judge tokens alongside accuracy because a memory system should improve the useful context available to an agent without treating an ever-larger prompt as the answer.

Every published score keeps its protocol, category breakdown, and policy boundary attached. We show genuine core runs; benchmark-only fallbacks and retrieval-only diagnostics do not become headline performance claims.

Confidence gate

We have not run the full 500-question suite. These diagnostics are the rehearsal: we will spend the full evaluation only when the system is stable enough that the result will teach us something reliable.

Diagnostic history

Progress across the frozen 12 questions

All points share the same question cohort and frozen retrieval envelope. Fresh answer and judge calls still introduce sampling variance.

50%60%70%80%8/12Run 17/12Run 28/12Run 37/12Run 49/12Run 5
RunOverallTemporalKnowledge updateTokens / question

Frozen raw top-six control

Control

Raw retrieved sessions supplied without the bounded compiler capsule. This is the fixed diagnostic control, not a product configuration.

8 / 12

66.67%

2 / 6

33.33%

6 / 6

100.00%

21,916

262,995 total

Compiler candidate 23dc6b7

Did not pass

Reduced prompt cost substantially, but regressed one answer against the frozen control.

7 / 12

58.33%

3 / 6

50.00%

4 / 6

66.67%

6,424

77,092 total

Compiler candidate ed70543

Did not pass

Matched the control overall and improved the temporal slice, but regressed knowledge-update questions.

8 / 12

66.67%

5 / 6

83.33%

3 / 6

50.00%

9,892

118,709 total

Compiler candidate eb1d397

Did not pass

Aggregate fallback narrowing did not restore the missing knowledge-update evidence.

7 / 12

58.33%

4 / 6

66.67%

3 / 6

50.00%

9,806

117,676 total

Temporal-win checkpoint 2b38545

Checkpoint

First accepted diagnostic uplift: one more correct answer than the control, with a much smaller context envelope.

9 / 12

75.00%

4 / 6

66.67%

5 / 6

83.33%

9,726

116,709 total

Scorecard

What we measured—and how to read it

75.00% · 9 / 12

Diagnostic accuracy

The best result on the frozen 12-question temporal and knowledge-update diagnostic.

What it is for

Use it to decide whether a compiler revision is ready for broader evaluation.

A 12-question diagnostic has high variance and must never be presented as the official 500-question LongMemEval result.

66.67% · 4 / 6

Temporal questions

Questions requiring dates, elapsed time, or changing state in the accepted checkpoint.

What it is for

Use it to expose temporal evidence and answer-assembly failures before a full run.

An earlier candidate reached 5/6 temporal but regressed knowledge-update performance, so slices are read together.

83.33% · 5 / 6

Knowledge updates

Questions requiring the system to retain and apply updated information.

What it is for

Use it to test whether compressed context preserves superseding facts.

Fresh stochastic answer and judge calls mean observed deltas are not isolated causal attributions.

9,726 tokens / question

Token efficiency

Total answer and judge tokens at the accepted checkpoint, versus 21,916 for the raw top-six control.

What it is for

Use it to track whether bounded evidence keeps evaluation cost under control.

Token totals cover this exact 12-question diagnostic protocol and should not be projected directly to production traffic.

Run protocol

Benchmark
LongMemEval fixed diagnostic
Question denominator
12 of the official suite's 500
Retrieval envelope
6 frozen retrieved sessions
Compiler budget
16 facts / 2,400 units
Answer model
GPT-5.6 Luna
Primary judge
GPT-5.6 Sol
Full-suite status
Not run; awaiting confidence gate

What this supports

  • Whether a context-compiler revision deserves a broader evaluation.
  • How temporal and knowledge-update slices trade off within the frozen diagnostic.
  • The token cost of raw retrieved sessions versus bounded compiler capsules.

What it does not support

  • An official or projected score on the 500-question LongMemEval suite.
  • A leaderboard comparison with systems evaluated on a different denominator or protocol.
  • A causal claim about one compiler change when answer and judge calls were freshly sampled.