A rubric isn't a scoring prompt — it's a five-step protocol that turns "the user complained" into "team X ships fix Y." Three guides, twenty-five questions, forty-seven atomic concepts, grounded in ten research sources.
All 25 questions in one sitting — live score tracks as you go, with a per-guide breakdown at the end.
Start quiz →47 flip cards across all guides. Filter by guide, shuffle the deck, or flip everything at once.
Browse cards →Why a total quality score can't drive a fix. The five-step minimum closed loop. The atomic dimension table that maps every failure to one asset layer. G-Eval's structured judgment. LLM Judge biases — position, verbosity, self-enhancement — with the actual MT-Bench numbers.
Read guide →Anthropic's grader stratification. Why deterministic checks must come before LLM-as-Judge. The three rubric kinds — Outcome, Answer, Trajectory — and which catches what. TRAJECT-Bench's four trajectory failure modes. Evidence prerequisites, UNKNOWN as first-class output, and pass@k vs pass^k for non-deterministic flows.
Why generic checklists pass dangerous solutions and fail valid ones. The contextual rubric chain — Context Gathering → Rubric Builder → Trace Inspector → Failure Attribution → Eval Case Generator. Agentic Rubrics for SWE Agents (2% flakiness with atomic items). The four rubric-aging signals. Treating rubric changes like system changes — bug report → patch → calibration replay → shadow eval → release.
Read guide →