CoachDiscoverMapCompareSavedAI LearningAI WeeklyAbout
A Self-Learning Library · For Product Managers, Jun 2026

The Agent Eval Rubric

A rubric isn't a scoring prompt — it's a five-step protocol that turns "the user complained" into "team X ships fix Y." Three guides, twenty-five questions, forty-seven atomic concepts, grounded in ten research sources.

Sources: G-Eval · MT-Bench · Anthropic Engineering · RAGAS · TRAJECT-Bench · Agentic Rubrics for SWE Agents
3
Learning Guides
47
Concepts Covered
25
Quiz Questions
10
Research Sources
Practice — test & drill what you've learned
?
Multiple Choice · All Guides
25Questions
3Guides
LiveScore

Full Quiz

All 25 questions in one sitting — live score tracks as you go, with a per-guide breakdown at the end.

Start quiz →
🃏
Flip Cards · All Guides
47Cards
3Guides
Shuffle+ Filter

All Flashcards

47 flip cards across all guides. Filter by guide, shuffle the deck, or flip everything at once.

Browse cards →
Curriculum — three guides, in learning order
01
01 · Foundations

Why Rubrics Are a Product DecisionFailure rarely lives where you think it does

Why a total quality score can't drive a fix. The five-step minimum closed loop. The atomic dimension table that maps every failure to one asset layer. G-Eval's structured judgment. LLM Judge biases — position, verbosity, self-enhancement — with the actual MT-Bench numbers.

Closed loopAtomic dimensionsG-EvalJudge biasesRAGAS split9 questions · 19 cards
Read guide
02
02 · Applied

Building Your RubricCode first, model second, human last

Anthropic's grader stratification. Why deterministic checks must come before LLM-as-Judge. The three rubric kinds — Outcome, Answer, Trajectory — and which catches what. TRAJECT-Bench's four trajectory failure modes. Evidence prerequisites, UNKNOWN as first-class output, and pass@k vs pass^k for non-deterministic flows.

Grader tiersOutcome / Answer / TrajectoryTRAJECT-BenchEvidence gatespass@k vs pass^k8 questions · 16 cards
Read guide
03
03 · Advanced

Operating the RubricRubrics age — treat them as system assets

Why generic checklists pass dangerous solutions and fail valid ones. The contextual rubric chain — Context Gathering → Rubric Builder → Trace Inspector → Failure Attribution → Eval Case Generator. Agentic Rubrics for SWE Agents (2% flakiness with atomic items). The four rubric-aging signals. Treating rubric changes like system changes — bug report → patch → calibration replay → shadow eval → release.

Contextual chainAgentic RubricsAging signalsShadow evaluationRubric versioning8 questions · 12 cards
Read guide
How to use this library.
  1. Read the guides in order — Foundations sets the language, Applied lets you build, Advanced lets you operate. Each guide is a 15–20 minute read.
  2. Drill with flashcards while reading — the deck has a filter for the current guide, so you can lock in the atoms as you go.
  3. Take the full quiz after all three guides. Many questions are PM-scenario judgments ("your team proposes X…") that test whether the principles transfer.
  4. Use the dimension table from Guide 01 as your reference card — print it, paste it in your team's eval doc, or keep it open during the next eval review.

Grounded in 10 research sources

  1. Yang Liu et al., G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, arXiv:2303.16634, 2023.
  2. Lianmin Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, arXiv:2306.05685, NeurIPS 2023.
  3. Anthropic Engineering, Demystifying evals for AI agents, anthropic.com/engineering.
  4. LangSmith Docs, How to define an LLM-as-a-judge evaluator.
  5. DeepEval Docs, G-Eval / The LLM Evaluation Framework.
  6. RAGAS Docs, Metrics.
  7. Google Cloud / Vertex AI Docs, Gen AI Evaluation Service API.
  8. Pengfei He et al., TRAJECT-Bench: A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use, arXiv:2510.04550, 2025.
  9. Runyang You et al., Agent-as-a-Judge survey, arXiv:2601.05111, 2026.
  10. Mohit Raghavendra et al., Agentic Rubrics as Contextual Verifiers for SWE Agents, arXiv:2601.04171, 2026.