The Eval Stack
Almost every AI feature I see comes with a number. Ninety-one percent accurate. Judge score up four points. Very few come with the answers that decide whether the number means anything: measured on what, graded by whom, and compared to what.
There is no such thing as “an eval” in the abstract. There are five independent choices: the unit being judged, where the data comes from, who or what does the grading, what shape the score takes, and when it runs. What gets sold as a framework is one combination of those, packaged. Two teams can both report 87% and be measuring unrelated things, and most arguments about evals are really a disagreement on one axis, argued as if it were about the whole thing.
This tool is for the person who decides what gets built, trusted and shipped, not the one writing the scorer. Start by describing your situation and get a concrete plan. Then take the plan apart: see what your combination can and cannot catch, watch a judge’s biases move a ship decision, and check whether your test set is big enough to support the claim you want to make.
Answer for the situation you are in, not the one you would like to be in. The plan names a dataset, a grader mix, a cadence, candidate tools, the trap specific to your case, and what to do in the first week.
- 150–300 real questions sampled from logs or from whoever answers these today, stratified by question type, plus every known hard case. Each needs a known-correct answer and the record it comes from.
- Split retrieval from generation; most ‘model’ failures are retrieval failures. Code checks that citations resolve and figures reconcile, a judge for groundedness and completeness, and a refusal check for questions the data cannot answer.
- Code assertions on every pull request; the full judge-graded suite on each release candidate; a one-time human review of 100 items to validate the judge before you trust it. Write the launch gates down now.
- A harness (Promptfoo or DeepEval) plus RAG-decomposed metrics (Ragas or TruLens), over traces in Phoenix, Langfuse or LangSmith.
- Measuring answer quality only. If you cannot see context precision and recall, you will tune prompts for months against a retrieval bug.
- Build the golden set from the failure taxonomy, then get 100 human labels so you have a κ for your judge before launch.
›Methods, and where this is deliberately simple
Sample sizes use the two-proportion normal approximation (unpaired) and Connor’s McNemar formula (paired), with critical values from an inverse-normal rather than a lookup table. The smallest detectable lift inverts the McNemar formula by bisection. The interval is Wilson’s score interval. The judge correction is Rogan–Gladen, and the shrinkage factor is Youden’s J (sensitivity + specificity − 1). These are exact for their stated assumptions. Everything runs in your browser and nothing you enter leaves the page.
The judge lab is a toy model, seeded so it draws the same figure every time: 40 items with fixed difficulty, Gaussian judge noise, and fixed bias terms for position, length and self-preference. The bias sizes are chosen to show direction, not measured from any real judge, and the κ it reports is a heuristic. The composer scores are coarse 0–5 ratings averaged across axes; read them as direction, not as a number to optimize. The correction assumes the judge’s error rates are known and stable, which is only true if you measured them on a labeled set.
The dataset is the eval
Teams spend most of their eval effort on graders and most of their eval risk on data. A clever judge on 40 questions an engineer wrote from memory tells you about those 40 questions. Sampled production traces are the only data that reflects what users actually do. Synthetic cases scale instantly and quietly share the blind spots of the model that wrote them. Adversarial cases are unrepresentative on purpose. A good set uses all of them, weighted on purpose, with a held-out slice nobody tunes against.
The unit matters as much. An agent can reach a perfect final answer after six wasteful tool calls and one dangerous write, and an eval that only reads the final answer scores it 1.0.
Five ways to produce a judgment
Every grader is a variation on five families. They differ by three orders of magnitude in cost and by a wide margin in how closely they track what users experience. A mature stack runs four of them at once, at different cadences.
| Family | Cost per item | Variance | Measures | How it fails |
|---|---|---|---|---|
| Programmatic checks | ~$0 | none | Form, verifiable facts | Passes confidently on useless answers |
| Reference metrics | ~$0 | none | Overlap with a gold answer | Penalizes correct paraphrase; rewards fluent wrongness |
| LLM judge | $0.002–0.05 | moderate | Stated criteria, open-ended | Position, verbosity, self-preference bias; silent drift |
| Human expert review | $1–10 | moderate | Quality, as the domain defines it | Low inter-rater agreement; too slow to gate on |
| Product signals | ~$0 | high | Realized user value | Confounded, sparse, weeks of lag |
The family that needs the most supervision is the one most teams lean on. An LLM judge scales to thousands of items and can be pointed at a new criterion in an afternoon, but it is a model with its own failure modes. Validate it against human labels before trusting it, and report chance-corrected agreement (Cohen’s κ), not raw percent agreement. A June 2026 study of 21 judge models found that exact-match agreement overstated performance by 33 to 41 points on MT-Bench once corrected for chance. It also found position bias spanning nearly two orders of magnitude across judges while test–retest reliability stayed above 0.95. A judge can be perfectly consistent and consistently wrong.
Human review sets the ceiling. If two of your own experts agree only 70% of the time, no automated grader validated against them can do better. The highest-value use of expert time is not scoring at volume. It is reading a few hundred real traces, naming the failure modes, and turning each into a cheap automated check.
What exists, and the trap in each
Named tools fall into five groups, and mixing them up is the most common mistake. A public benchmark tells you something about a model. A harness runs your evals. An observability platform captures production traces to evaluate. A metric library supplies scorers. A red-team suite probes for harm. Most teams need one from each of three groups, and none from the first.
MMLU / MMLU-Pro
Broad multiple-choice knowledge across 57 subjects; MMLU-Pro is the harder, 10-option revision.
MMLU / MMLU-Pro
Saturated at the frontier and widely contaminated by training data. Useful for ruling out weak models, close to useless for ranking strong ones, and silent about your application.
GPQA Diamond
Graduate-level physics, chemistry and biology questions written to be Google-proof.
GPQA Diamond
A reasoning-depth signal only. The set is small, so published scores carry wide intervals that vendors rarely show.
SWE-bench Verified / Pro
Real GitHub issues; the model’s patch must make the repo’s own tests pass. Pro is the harder, less contaminated successor.
SWE-bench Verified / Pro
In February 2026 OpenAI stopped reporting Verified, after an audit found flawed tests in most of the problems it checked and evidence that frontier models had seen the fixes in training. Scaffold and retry budget also move scores, so ask what harness produced any figure.
Terminal-Bench 2
End-to-end terminal tasks in a container: install, configure, debug, verify. Graded on final state.
Terminal-Bench 2
Environment-sensitive, and versions are not comparable with each other. A good proxy for agentic competence on real machines, a weak one for your domain tooling.
τ-bench family (τ², τ³)
Tool-using agents in business domains against a simulated user, with database-state verification and pass^k for consistency. τ³ adds voice and knowledge-retrieval tasks.
τ-bench family (τ², τ³)
The most relevant public shape for business agents, and the clearest proof that a good pass@1 can hide terrible reliability. The domains are not yours; borrow the method, not the score.
GAIA
Assistant tasks that need browsing, file handling and multi-step reasoning, with short verifiable answers.
GAIA
Answer-matching only: it grades the destination, not whether the agent took a sane or safe route to get there.
WebArena / OSWorld
Self-hosted web and desktop environments for computer-use agents, graded on achieved state.
WebArena / OSWorld
Heavy to run and sensitive to environment versions. Capability research, not a routine gate.
MTEB
Massive Text Embedding Benchmark: retrieval, clustering, classification and reranking across many datasets.
MTEB
The practical way to shortlist embedding models, but leaderboard rank rarely survives contact with your corpus. Re-run the retrieval tasks on your own documents.
Arena (formerly LMArena)
Human pairwise preference at scale, aggregated into model ratings, with style-controlled views.
Arena (formerly LMArena)
Measures what anonymous users prefer, which rewards formatting, length and friendliness. Style control helps. Preference is not correctness.
HELM
Stanford's standardized harness reporting many metrics per scenario: accuracy, calibration, robustness, fairness, efficiency.
HELM
The multi-metric discipline is the valuable part and worth copying. As a leaderboard it lags the frontier.
LiveBench
Contamination-resistant benchmark that refreshes questions from recent sources on a rolling basis.
LiveBench
Freshness is the point, which also means scores are not comparable across refreshes. Read the trend, not the absolute.
ARC-AGI-2 / ARC-AGI-3
Abstract reasoning puzzles built to resist memorization. ARC-AGI-3 moves to interactive environments an agent has to explore without instructions.
ARC-AGI-2 / ARC-AGI-3
A research signal about generalization, with no bearing on whether a model can handle your workflow.
lm-evaluation-harness
EleutherAI's standard runner for academic benchmarks; what most published model numbers are produced with.
lm-evaluation-harness
Built for comparing models on static datasets, not for evaluating an application. Prompt-format choices move scores by several points, so pin your configuration.
OpenAI Evals
An open-source registry and runner (declare a dataset plus a grader), and a hosted Evals API on the OpenAI platform.
OpenAI Evals
The open-source runner is minimal; you build reporting, dataset management and CI wiring yourself. The hosted API is convenient and ties results to one vendor.
Promptfoo
Declarative test cases with assertions, matrix comparison across prompts and models, and a red-team generator. Part of OpenAI since March 2026, still MIT-licensed open source.
Promptfoo
The fastest route from zero to a blocking CI gate. Large suites in YAML get unwieldy, so move complex graders into code before the config becomes the problem.
DeepEval
Pytest-style unit testing for LLM apps, with bundled metrics including G-Eval-style judges, hallucination and relevancy scorers.
DeepEval
Feels natural to engineers because it is just tests. Treat the bundled metrics as drafts and validate any judge against your own labels before trusting a threshold.
Inspect AI
The UK AI Security Institute's framework: datasets, solvers, scorers and tools as composable pieces, with strong logging and provenance. Built for agentic and safety evaluation.
Inspect AI
The most rigorous open option, and the right choice when an eval has to be defensible or audited. A steeper learning curve than the YAML tools.
MLflow GenAI evaluation
mlflow.genai.evaluate with built-in and custom LLM-judge scorers, inside an existing MLflow tracking setup.
MLflow GenAI evaluation
Pick it for continuity with an established MLOps or Databricks stack, not for LLM-specific depth.
Microsoft Foundry / Gemini Enterprise Agent Platform
Cloud-native evaluation inside the Azure and Google agent platforms (formerly Azure AI Studio and Vertex AI), with built-in judges, safety evaluators and managed runs.
Microsoft Foundry / Gemini Enterprise Agent Platform
Lowest friction if you are already committed to that cloud and need auditability. Portability is the cost: judge definitions and results do not travel.
Ragas
RAG-specific metrics: faithfulness, answer relevancy, context precision and context recall, per sample.
Ragas
The decomposition is the value: it separates ‘retrieval missed it’ from ‘the model ignored it’. Several metrics are themselves judge-based and inherit judge error. Context recall needs ground-truth references.
TruLens
Feedback functions over traced app runs, including the RAG triad: context relevance, groundedness, answer relevance.
TruLens
A good instrumentation-first model. The triad is a diagnostic frame, not a release gate on its own.
G-Eval (a method, not a tool)
Judge pattern: generate evaluation steps from a criterion, then score with chain-of-thought and token-probability weighting.
G-Eval (a method, not a tool)
Correlates better with human judgment than a bare ‘rate 1–10’ prompt, and several libraries implement it. Still a judge, so validate it.
LangSmith
Tracing, dataset curation, experiments and annotation queues, with native LangChain and LangGraph integration plus OpenTelemetry.
LangSmith
Strongest when you already use that ecosystem. The annotation queue is underrated: it is how production traces become a golden set.
Braintrust
An eval-driven development loop: prompt playground, versioned experiments, scorers, and production logs feeding back into datasets.
Braintrust
Built around iterating on prompts with evidence. Commercial, with self-hosted data-plane options.
Arize Phoenix / Arize AX
OpenTelemetry-native tracing and evaluation at span, trace, trajectory and session level. Phoenix is the free, self-hostable version.
Arize Phoenix / Arize AX
The OTel schema keeps your traces portable across vendors, a real consideration when choosing. Agent trajectory evaluation is a genuine strength.
Langfuse
MIT-licensed core for tracing, prompt management, datasets and evaluation rules, with OTel-compatible SDKs. Part of ClickHouse since January 2026.
Langfuse
The common choice when self-hosting is a requirement. Evaluation features are lighter than the dedicated eval platforms, so many teams pair it with a harness.
W&B Weave
Tracing of agent sessions, turns and tools, with custom scorers, monitors and production guardrails.
W&B Weave
Natural if your team already lives in Weights & Biases. Agent-level evaluation leans on scorers you write.
Comet Opik
Apache-2.0 tracing and evaluation with step- and thread-level agent scoring and an OTLP endpoint.
Comet Opik
Permissive license and agent-oriented. A smaller ecosystem than the leaders, so check integration coverage for your stack.
HarmBench / JailbreakBench
Standardized sets of harmful behaviors and jailbreak attempts with defined attack and refusal protocols.
HarmBench / JailbreakBench
Comparable numbers across systems, which is their purpose. They do not cover your domain's specific misuse, so write your own adversarial set as well.
AgentHarm
Whether tool-using agents refuse harmful multi-step tasks, and whether refusal survives jailbreaking.
AgentHarm
Matters because agent harm compounds: a model that says something bad costs less than one that acts on it.
garak
Open-source vulnerability scanner for LLMs: probes for prompt injection, leakage, toxicity, encoding attacks and more.
garak
Scanner-style breadth, so expect false positives and triage. Good as a scheduled sweep, poor as a hard blocking gate.
PyRIT
Microsoft's automation framework for AI red-teaming: orchestrators, attack strategies and scoring for repeatable campaigns.
PyRIT
A framework, not a test set; it needs someone who knows what to probe for. Pairs well with a standardized benchmark for the baseline.
Build it in this order
Eval maturity is a sequence, not a menu. Teams that jump to the third rung before doing the first build a precise, automated measurement of the wrong thing.
- 00
Vibes, plus logging
Someone tries ten prompts by hand and forms an opinion. Legitimate during a prototype, because you do not yet know what failure looks like. The one non-negotiable is logging every input, output, tool call and retrieved chunk from day one, with a trace ID. Everything above this rung is built from those logs.
- 01
Error analysis and a failure taxonomy
Read a hundred or two real traces, label each failure in your own words, and cluster the notes into named failure modes with counts. It is the highest-value activity in the discipline and the one most often skipped, because it is manual reading and cannot be bought.
- 02
A golden set with programmatic assertions
Turn the top failure modes into code. Build 100–500 cases weighted by real frequency, deliberately including the hard tail, and write deterministic checks. Run it on every change, and pin a held-out slice you never look at so you can tell overfitting from improvement.
- 03
A validated LLM judge in CI
Write a rubric of specific yes/no criteria, not “rate helpfulness 1–10.” Have people label 100–200 items, measure the judge’s agreement as Cohen’s κ, and iterate until it clears about 0.6 for most product decisions, higher when stakes are. Pin the judge version and report intervals, never bare point estimates.
- 04
Online evaluation and monitoring
Canary or A/B the change on live traffic with guardrails declared in advance: quality, cost per task, p95 latency, refusal and escalation rate, safety incidents. Sample production through the same judge, segmented by cohort. Feed new failures back into the golden set; a set that never changes is one you have overfit.
Notice what this implies about staffing. The first two rungs are careful human work on data, not model work. If an eval plan has no line item for someone reading traces, it is not a plan.
How evals lie
Each of these produces a number that is technically correct and practically misleading. They are listed roughly by how often they cost real money.
The dataset nobody questions
Sixty cases an engineer wrote from memory in the first sprint, still the release gate a year later. It measures the team’s early assumptions, faithfully, forever. Counter: Re-derive the set from sampled production traces every quarter, and keep a held-out slice you never tune against.
Point estimates with no interval
“87% vs 85%” on 200 items is almost always noise, and it gets reported as progress because nobody computed the interval. Counter: Report every number with an interval and a sample size, and pair the comparison so the interval is tight enough to mean something.
The unvalidated judge
A judge prompt written in an afternoon, never compared to human labels, producing a number quoted in planning for a year. It is consistent, which reads as reliable. Consistency is not accuracy. Counter: Measure Cohen’s κ against human labels before first use, and again whenever the judge model, rubric or product changes.
Goodharting the judge
Once a judge score becomes a target, prompts get tuned to the judge’s quirks. The score climbs; users notice nothing. Counter: Keep a human-labeled holdout the optimization loop never sees, and watch product signals for divergence.
Averages that hide regressions
One headline number stays flat while the segment your biggest customer uses falls eight points. Counter: Name the segments that matter before the test and gate on the worst one, not the mean.
Benchmark scores as product evidence
A public leaderboard decides a model choice, and the number came from a different scaffold, prompt format and retry budget than yours. Counter: Use public benchmarks to shortlist; decide on your own eval set, with every candidate run through your harness.
Contamination and saturation
Popular benchmarks leak into training data, and the frontier clusters within a point where there is no headroom left. Counter: Prefer held-out, refreshing or private sets. Treat any score in the mid-nineties as information about the benchmark, not the model.
Testing the prompt, not the system
Prompt variants get compared for weeks while the real cause is chunking, retrieval ranking, a stale index or a tool timeout. Counter: Evaluate components separately. Retrieval precision and recall, tool success rate and end-to-end quality are three different measurements.
Quality-only scorecards
A change ships on a 3-point quality gain and doubles cost per task and p95 latency. Both were measured; neither was on the gate. Counter: Put cost per resolved task, p95 latency, refusal and escalation rate, and safety incidents on the same scorecard as quality.
Silent judge drift
The judge’s underlying model is upgraded, every historical number shifts, and the trend line is quietly meaningless. Counter: Pin judge versions, re-run the baseline whenever one changes, and store the judge version with every result.
Questions to ask before you trust the number
You do not need to write a scorer to hold an eval to a standard. These separate a measurement from a vibe, and none of them require reading the code.
- 01
What is in the eval set, who built it, and how do we know it looks like real traffic?
The honest answer is often “an engineer wrote 40 cases from memory.” That is a smoke test. Ask what share came from sampled production traces.
- 02
What is the confidence interval on that number?
If nobody has one, the comparison probably cannot support the decision. The Size step above does the arithmetic.
- 03
Is this judged by code, a model or a person, and if a model, what is its agreement with human labels?
Ask for Cohen’s κ, not percent agreement, which overstates badly when one class dominates.
- 04
Were both systems run on the same items?
Paired comparisons need a fraction of the data. An unpaired comparison on a small set is usually inconclusive.
- 05
Which segments got worse?
A flat average hides a regression in the cohort that complains loudest. Ask for the breakdown before the headline.
- 06
What is the p95, not the mean?
Users experience the tail. This applies to quality as much as latency.
- 07
What does each task cost, and how long does it take, at this quality?
Quality is one axis of four: quality, cost, latency, safety. A scorecard with only quality on it is incomplete.
- 08
When did the judge model or prompt last change, and was the baseline re-run?
An unannounced judge upgrade can shift every historical number.
- 09
What did the last error analysis find, and when was it done?
If the answer is “we haven’t,” the suite is testing assumptions rather than reality.
- 10
Would this suite have caught the failure that got escalated last week?
Usually it would not, and that gap is your next sprint.
A scorecard that can carry a ship decision
The figures are illustrative. The headline would have read “quality up 2.8 points.”
| Metric | Grader | Baseline | Candidate | Δ (95% CI) | Gate |
|---|---|---|---|---|---|
| Output schema valid | code | 99.4% | 99.8% | +0.4 (−0.1, +0.9) | |
| Cited record exists | code | 96.1% | 98.7% | +2.6 (+1.4, +3.8) | |
| Answer grounded in data | judge, κ 0.71 | 88.2% | 91.0% | +2.8 (+0.6, +5.0) | |
| Correct line item identified | judge, κ 0.68 | 84.5% | 84.1% | −0.4 (−3.1, +2.3) | |
| Multi-line POs only | judge, κ 0.68 | 79.0% | 71.2% | −7.8 (−12.9, −2.7) | regressed |
| Refuses when data missing | code + judge | 91.0% | 92.4% | +1.4 (−1.0, +3.8) | |
| Cost per resolved question | telemetry | $0.031 | $0.052 | +68% | review |
| p95 latency | telemetry | 4.1 s | 7.8 s | +90% | breach |
The rows that matter are the segment that fell 7.8 points and the latency breach. Both are invisible in an average, and both decide the launch. A scorecard’s job is to make those two rows impossible to miss.
If you do one thing
Read a hundred real traces this week and write down what went wrong, in your own words, with counts. Everything above is downstream of that list. The tools get chosen in an afternoon. The taxonomy is what takes judgment, and it is the part nobody else can do for you.