We evaluate Copepod as a complete system: a modern answer model working with a durable memory layer. We are not trying to reproduce another memory company's architecture or tune our work to match a competitor's reported score. We use capable modern models deliberately, build the memory system we believe agents need, and measure whether it improves over time.
Quality and efficiency move together. We report answer and judge tokens alongside accuracy because a memory system should improve the useful context available to an agent without treating an ever-larger prompt as the answer.
Every published score keeps its protocol, category breakdown, and policy boundary attached. We show genuine core runs; benchmark-only fallbacks and retrieval-only diagnostics do not become headline performance claims.