Copepod research
Results you can inspect, not just repeat.
Every benchmark report defines what was measured, the protocol behind it, and the boundary of the claim. We publish retrieval, answer quality, evaluator sensitivity, and latency as distinct measurements when available.
Published reports
Benchmark white papers
Measuring long-term conversational memory on LoCoMo
A transparent end-to-end evaluation of Copepod's ability to retrieve and answer questions from long, multi-session conversations.
Building confidence on LongMemEval before the full suite
A transparent development timeline on a frozen 12-question diagnostic set, used to improve temporal and knowledge-update memory before spending a full 500-question evaluation.
Full protocol
Dataset version, question denominator, system variant, model roles, and measurement boundary.
Claim boundaries
What each number supports—and what it cannot establish—written beside every score.
Comparable evidence
Diagnostic slices stay separate from accepted end-to-end results, so comparisons remain honest.