CoachDiscoverMapCompareSavedAI LearningAI WeeklyAbout
A Learning Library · Anthropic Engineering, Jan 2026

Demystifying Evals for AI Agents

Three guides distilling Anthropic's playbook for building eval suites — from "what's an eval" all the way to a zero-to-one roadmap for shipping agents with confidence.

Source: anthropic.com/engineering · Authors: Grace · Hadfield · Olivares · De Jonghe
3
Learning Guides
40+
Concepts Covered
40
Flashcards
8
Roadmap Steps
Practice — drill the vocabulary
🃏
Flip Cards · All Guides
40Cards
3Guides
Shuffle+ Filter

All Flashcards

Forty flip cards covering every concept across the three guides. Filter by guide, shuffle the deck, or flip everything at once.

Browse cards →
Curriculum — three guides, in learning order
01
01 · Foundations

The Anatomy of an EvalWhat it is, what to call its parts, and how graders work

Start here. Build the vocabulary you'll need to read everything else: task, trial, grader, transcript, outcome, harness, suite. Then meet the three families of graders — code-based, model-based, human — and see why capability and regression evals are different jobs.

DefinitionsGradersSingle vs multi-turnCapability vs regression17 cards
Read guide
02
02 · Applied

Four Agent Types, One PatternCoding, conversational, research, computer-use — and non-determinism

How evaluation actually looks for the four common agent shapes, including the benchmarks that matter (SWE-bench Verified, τ2-Bench, BrowseComp, WebArena, OSWorld). Closes with pass@k and pass^k — why a 75% pass-rate agent fails for customer-facing reliability.

Coding agentsConversationalResearchComputer usepass@k / pass^k11 cards
Read guide
03
03 · Advanced

Zero to One: A RoadmapThe eight-step playbook plus Swiss-cheese coverage

The full eight-step roadmap: collect tasks from real failures, design a stable harness, sweat the grader details, read the transcripts, watch for saturation, and treat eval suites as living artifacts. Closes with the Swiss-cheese model of how automated evals stack with monitoring, A/B tests, and human review.

RoadmapEval-driven devSaturationTranscript reviewSwiss cheeseFrameworks12 cards
Read guide
How to use this library. Read Guide 01 cover-to-cover for the vocabulary, skim Guide 02 for the agent type you're building, and treat Guide 03 as a reference you come back to when planning a real suite. Drill the flashcards anytime you want to refresh terminology — the deck is shuffled across all three guides.