Three guides distilling Anthropic's playbook for building eval suites — from "what's an eval" all the way to a zero-to-one roadmap for shipping agents with confidence.
Forty flip cards covering every concept across the three guides. Filter by guide, shuffle the deck, or flip everything at once.
Browse cards →Start here. Build the vocabulary you'll need to read everything else: task, trial, grader, transcript, outcome, harness, suite. Then meet the three families of graders — code-based, model-based, human — and see why capability and regression evals are different jobs.
Read guide →How evaluation actually looks for the four common agent shapes, including the benchmarks that matter (SWE-bench Verified, τ2-Bench, BrowseComp, WebArena, OSWorld). Closes with pass@k and pass^k — why a 75% pass-rate agent fails for customer-facing reliability.
Read guide →The full eight-step roadmap: collect tasks from real failures, design a stable harness, sweat the grader details, read the transcripts, watch for saturation, and treat eval suites as living artifacts. Closes with the Swiss-cheese model of how automated evals stack with monitoring, A/B tests, and human review.
Read guide →