Copepod research

Results you can inspect, not just repeat.

Every benchmark report defines what was measured, the protocol behind it, and the boundary of the claim. We publish retrieval, answer quality, evaluator sensitivity, and latency as distinct measurements when available.

Published reports

Benchmark white papers

Full protocol

Dataset version, question denominator, system variant, model roles, and measurement boundary.

Claim boundaries

What each number supports—and what it cannot establish—written beside every score.

Comparable evidence

Diagnostic slices stay separate from accepted end-to-end results, so comparisons remain honest.