Tools

The Eval Stack

Almost every AI feature I see comes with a number. Ninety-one percent accurate. Judge score up four points. Very few come with the answers that decide whether the number means anything: measured on what, graded by whom, and compared to what.

There is no such thing as “an eval” in the abstract. There are five independent choices: the unit being judged, where the data comes from, who or what does the grading, what shape the score takes, and when it runs. What gets sold as a framework is one combination of those, packaged. Two teams can both report 87% and be measuring unrelated things, and most arguments about evals are really a disagreement on one axis, argued as if it were about the whole thing.

This tool is for the person who decides what gets built, trusted and shipped, not the one writing the scorer. Start by describing your situation and get a concrete plan. Then take the plan apart: see what your combination can and cannot catch, watch a judge’s biases move a ship decision, and check whether your test set is big enough to support the claim you want to make.

Answer for the situation you are in, not the one you would like to be in. The plan names a dataset, a grader mix, a cadence, candidate tools, the trap specific to your case, and what to do in the first week.

What are you shipping?
Where is it?
What happens when it is wrong?
What labeling capacity do you have?

Your stack

Dataset
150–300 real questions sampled from logs or from whoever answers these today, stratified by question type, plus every known hard case. Each needs a known-correct answer and the record it comes from.
Graders
Split retrieval from generation; most ‘model’ failures are retrieval failures. Code checks that citations resolve and figures reconcile, a judge for groundedness and completeness, and a refusal check for questions the data cannot answer.
Cadence
Code assertions on every pull request; the full judge-graded suite on each release candidate; a one-time human review of 100 items to validate the judge before you trust it. Write the launch gates down now.
Tooling
A harness (Promptfoo or DeepEval) plus RAG-decomposed metrics (Ragas or TruLens), over traces in Phoenix, Langfuse or LangSmith.
The trap
Measuring answer quality only. If you cannot see context precision and recall, you will tune prompts for months against a retrieval bug.
First week
Build the golden set from the failure taxonomy, then get 100 human labels so you have a κ for your judge before launch.
›Methods, and where this is deliberately simple

Sample sizes use the two-proportion normal approximation (unpaired) and Connor’s McNemar formula (paired), with critical values from an inverse-normal rather than a lookup table. The smallest detectable lift inverts the McNemar formula by bisection. The interval is Wilson’s score interval. The judge correction is Rogan–Gladen, and the shrinkage factor is Youden’s J (sensitivity + specificity − 1). These are exact for their stated assumptions. Everything runs in your browser and nothing you enter leaves the page.

The judge lab is a toy model, seeded so it draws the same figure every time: 40 items with fixed difficulty, Gaussian judge noise, and fixed bias terms for position, length and self-preference. The bias sizes are chosen to show direction, not measured from any real judge, and the κ it reports is a heuristic. The composer scores are coarse 0–5 ratings averaged across axes; read them as direction, not as a number to optimize. The correction assumes the judge’s error rates are known and stable, which is only true if you measured them on a labeled set.

The dataset is the eval

Teams spend most of their eval effort on graders and most of their eval risk on data. A clever judge on 40 questions an engineer wrote from memory tells you about those 40 questions. Sampled production traces are the only data that reflects what users actually do. Synthetic cases scale instantly and quietly share the blind spots of the model that wrote them. Adversarial cases are unrepresentative on purpose. A good set uses all of them, weighted on purpose, with a held-out slice nobody tunes against.

The unit matters as much. An agent can reach a perfect final answer after six wasteful tool calls and one dangerous write, and an eval that only reads the final answer scores it 1.0.

Five ways to produce a judgment

Every grader is a variation on five families. They differ by three orders of magnitude in cost and by a wide margin in how closely they track what users experience. A mature stack runs four of them at once, at different cadences.

0255075100Closeness to real user experience ↑$0.0001$0.001$0.01$0.10$1$5Cost per judgment →Programmatic checks: ~$0 per judgment, answers in secondsProgrammatic checkssecondsReference metrics: ~$0 per judgment, answers in secondsReference metricssecondsLLM judge, single: $0.002–0.05 per judgment, answers in minutesLLM judge, singleminutesJudge jury + rubric: $0.01–0.15 per judgment, answers in minutesJudge jury + rubricminutesHuman expert review: $1–10 per judgment, answers in daysHuman expert reviewdaysProduct signals: ~$0 per judgment, answers in weeksProduct signalsweeks
Cost per judgment (log scale) against how closely the judgment tracks what users experience. Bigger dots take longer to answer: seconds for code, weeks for product signals. Product signals sit in the corner you want, cheap and close to the truth, but they arrive weeks late and are hard to pin on any one change. Everything faster gives up fidelity or costs more.
FamilyCost per itemVarianceMeasuresHow it fails
Programmatic checks~$0noneForm, verifiable factsPasses confidently on useless answers
Reference metrics~$0noneOverlap with a gold answerPenalizes correct paraphrase; rewards fluent wrongness
LLM judge$0.002–0.05moderateStated criteria, open-endedPosition, verbosity, self-preference bias; silent drift
Human expert review$1–10moderateQuality, as the domain defines itLow inter-rater agreement; too slow to gate on
Product signals~$0highRealized user valueConfounded, sparse, weeks of lag

Costs are order-of-magnitude figures for planning, not quotes.

The family that needs the most supervision is the one most teams lean on. An LLM judge scales to thousands of items and can be pointed at a new criterion in an afternoon, but it is a model with its own failure modes. Validate it against human labels before trusting it, and report chance-corrected agreement (Cohen’s κ), not raw percent agreement. A June 2026 study of 21 judge models found that exact-match agreement overstated performance by 33 to 41 points on MT-Bench once corrected for chance. It also found position bias spanning nearly two orders of magnitude across judges while test–retest reliability stayed above 0.95. A judge can be perfectly consistent and consistently wrong.

Human review sets the ceiling. If two of your own experts agree only 70% of the time, no automated grader validated against them can do better. The highest-value use of expert time is not scoring at volume. It is reading a few hundred real traces, naming the failure modes, and turning each into a cheap automated check.

What exists, and the trap in each

Named tools fall into five groups, and mixing them up is the most common mistake. A public benchmark tells you something about a model. A harness runs your evals. An observability platform captures production traces to evaluate. A metric library supplies scorers. A red-team suite probes for harm. Most teams need one from each of three groups, and none from the first.

Kind
Useful at
32 of 32 shown

MMLU / MMLU-Pro

Benchmark

Broad multiple-choice knowledge across 57 subjects; MMLU-Pro is the harder, 10-option revision.

›The trap

Saturated at the frontier and widely contaminated by training data. Useful for ruling out weak models, close to useless for ranking strong ones, and silent about your application.

Useful at: Model selection

GPQA Diamond

Benchmark

Graduate-level physics, chemistry and biology questions written to be Google-proof.

›The trap

A reasoning-depth signal only. The set is small, so published scores carry wide intervals that vendors rarely show.

Useful at: Model selection

SWE-bench Verified / Pro

Benchmark

Real GitHub issues; the model’s patch must make the repo’s own tests pass. Pro is the harder, less contaminated successor.

›The trap

In February 2026 OpenAI stopped reporting Verified, after an audit found flawed tests in most of the problems it checked and evidence that frontier models had seen the fixes in training. Scaffold and retry budget also move scores, so ask what harness produced any figure.

Useful at: Model selection

Terminal-Bench 2

Benchmark

End-to-end terminal tasks in a container: install, configure, debug, verify. Graded on final state.

›The trap

Environment-sensitive, and versions are not comparable with each other. A good proxy for agentic competence on real machines, a weak one for your domain tooling.

Useful at: Model selection

τ-bench family (τ², τ³)

Benchmark

Tool-using agents in business domains against a simulated user, with database-state verification and pass^k for consistency. τ³ adds voice and knowledge-retrieval tasks.

›The trap

The most relevant public shape for business agents, and the clearest proof that a good pass@1 can hide terrible reliability. The domains are not yours; borrow the method, not the score.

Useful at: Model selection

GAIA

Benchmark

Assistant tasks that need browsing, file handling and multi-step reasoning, with short verifiable answers.

›The trap

Answer-matching only: it grades the destination, not whether the agent took a sane or safe route to get there.

Useful at: Model selection

WebArena / OSWorld

Benchmark

Self-hosted web and desktop environments for computer-use agents, graded on achieved state.

›The trap

Heavy to run and sensitive to environment versions. Capability research, not a routine gate.

Useful at: Model selection

MTEB

Benchmark

Massive Text Embedding Benchmark: retrieval, clustering, classification and reranking across many datasets.

›The trap

The practical way to shortlist embedding models, but leaderboard rank rarely survives contact with your corpus. Re-run the retrieval tasks on your own documents.

Useful at: Model selection · Development

Arena (formerly LMArena)

Benchmark

Human pairwise preference at scale, aggregated into model ratings, with style-controlled views.

›The trap

Measures what anonymous users prefer, which rewards formatting, length and friendliness. Style control helps. Preference is not correctness.

Useful at: Model selection

HELM

Benchmark

Stanford's standardized harness reporting many metrics per scenario: accuracy, calibration, robustness, fairness, efficiency.

›The trap

The multi-metric discipline is the valuable part and worth copying. As a leaderboard it lags the frontier.

Useful at: Model selection

LiveBench

Benchmark

Contamination-resistant benchmark that refreshes questions from recent sources on a rolling basis.

›The trap

Freshness is the point, which also means scores are not comparable across refreshes. Read the trend, not the absolute.

Useful at: Model selection

ARC-AGI-2 / ARC-AGI-3

Benchmark

Abstract reasoning puzzles built to resist memorization. ARC-AGI-3 moves to interactive environments an agent has to explore without instructions.

›The trap

A research signal about generalization, with no bearing on whether a model can handle your workflow.

Useful at: Model selection

lm-evaluation-harness

Harness

EleutherAI's standard runner for academic benchmarks; what most published model numbers are produced with.

›The trap

Built for comparing models on static datasets, not for evaluating an application. Prompt-format choices move scores by several points, so pin your configuration.

Useful at: Model selection

OpenAI Evals

Harness

An open-source registry and runner (declare a dataset plus a grader), and a hosted Evals API on the OpenAI platform.

›The trap

The open-source runner is minimal; you build reporting, dataset management and CI wiring yourself. The hosted API is convenient and ties results to one vendor.

Useful at: Development · CI gate

Promptfoo

Harness

Declarative test cases with assertions, matrix comparison across prompts and models, and a red-team generator. Part of OpenAI since March 2026, still MIT-licensed open source.

›The trap

The fastest route from zero to a blocking CI gate. Large suites in YAML get unwieldy, so move complex graders into code before the config becomes the problem.

Useful at: Development · CI gate · Release

DeepEval

Harness

Pytest-style unit testing for LLM apps, with bundled metrics including G-Eval-style judges, hallucination and relevancy scorers.

›The trap

Feels natural to engineers because it is just tests. Treat the bundled metrics as drafts and validate any judge against your own labels before trusting a threshold.

Useful at: Development · CI gate

Inspect AI

Harness

The UK AI Security Institute's framework: datasets, solvers, scorers and tools as composable pieces, with strong logging and provenance. Built for agentic and safety evaluation.

›The trap

The most rigorous open option, and the right choice when an eval has to be defensible or audited. A steeper learning curve than the YAML tools.

Useful at: Development · CI gate · Release

MLflow GenAI evaluation

Harness

mlflow.genai.evaluate with built-in and custom LLM-judge scorers, inside an existing MLflow tracking setup.

›The trap

Pick it for continuity with an established MLOps or Databricks stack, not for LLM-specific depth.

Useful at: Development · CI gate · Production

Microsoft Foundry / Gemini Enterprise Agent Platform

Harness

Cloud-native evaluation inside the Azure and Google agent platforms (formerly Azure AI Studio and Vertex AI), with built-in judges, safety evaluators and managed runs.

›The trap

Lowest friction if you are already committed to that cloud and need auditability. Portability is the cost: judge definitions and results do not travel.

Useful at: Development · Release · Production

Ragas

Metrics

RAG-specific metrics: faithfulness, answer relevancy, context precision and context recall, per sample.

›The trap

The decomposition is the value: it separates ‘retrieval missed it’ from ‘the model ignored it’. Several metrics are themselves judge-based and inherit judge error. Context recall needs ground-truth references.

Useful at: Development · CI gate

TruLens

Metrics

Feedback functions over traced app runs, including the RAG triad: context relevance, groundedness, answer relevance.

›The trap

A good instrumentation-first model. The triad is a diagnostic frame, not a release gate on its own.

Useful at: Development · Production

G-Eval (a method, not a tool)

Metrics

Judge pattern: generate evaluation steps from a criterion, then score with chain-of-thought and token-probability weighting.

›The trap

Correlates better with human judgment than a bare ‘rate 1–10’ prompt, and several libraries implement it. Still a judge, so validate it.

Useful at: Development

LangSmith

Platform

Tracing, dataset curation, experiments and annotation queues, with native LangChain and LangGraph integration plus OpenTelemetry.

›The trap

Strongest when you already use that ecosystem. The annotation queue is underrated: it is how production traces become a golden set.

Useful at: Development · CI gate · Production

Braintrust

Platform

An eval-driven development loop: prompt playground, versioned experiments, scorers, and production logs feeding back into datasets.

›The trap

Built around iterating on prompts with evidence. Commercial, with self-hosted data-plane options.

Useful at: Development · CI gate · Production

Arize Phoenix / Arize AX

Platform

OpenTelemetry-native tracing and evaluation at span, trace, trajectory and session level. Phoenix is the free, self-hostable version.

›The trap

The OTel schema keeps your traces portable across vendors, a real consideration when choosing. Agent trajectory evaluation is a genuine strength.

Useful at: Development · Production

Langfuse

Platform

MIT-licensed core for tracing, prompt management, datasets and evaluation rules, with OTel-compatible SDKs. Part of ClickHouse since January 2026.

›The trap

The common choice when self-hosting is a requirement. Evaluation features are lighter than the dedicated eval platforms, so many teams pair it with a harness.

Useful at: Development · CI gate · Production

W&B Weave

Platform

Tracing of agent sessions, turns and tools, with custom scorers, monitors and production guardrails.

›The trap

Natural if your team already lives in Weights & Biases. Agent-level evaluation leans on scorers you write.

Useful at: Development · Production

Comet Opik

Platform

Apache-2.0 tracing and evaluation with step- and thread-level agent scoring and an OTLP endpoint.

›The trap

Permissive license and agent-oriented. A smaller ecosystem than the leaders, so check integration coverage for your stack.

Useful at: Development · CI gate · Production

HarmBench / JailbreakBench

Red team

Standardized sets of harmful behaviors and jailbreak attempts with defined attack and refusal protocols.

›The trap

Comparable numbers across systems, which is their purpose. They do not cover your domain's specific misuse, so write your own adversarial set as well.

Useful at: Release

AgentHarm

Red team

Whether tool-using agents refuse harmful multi-step tasks, and whether refusal survives jailbreaking.

›The trap

Matters because agent harm compounds: a model that says something bad costs less than one that acts on it.

Useful at: Release

garak

Red team

Open-source vulnerability scanner for LLMs: probes for prompt injection, leakage, toxicity, encoding attacks and more.

›The trap

Scanner-style breadth, so expect false positives and triage. Good as a scheduled sweep, poor as a hard blocking gate.

Useful at: CI gate · Release

PyRIT

Red team

Microsoft's automation framework for AI red-teaming: orchestrators, attack strategies and scoring for repeatable campaigns.

›The trap

A framework, not a test set; it needs someone who knows what to probe for. Pairs well with a standardized benchmark for the baseline.

Useful at: Release

32 entries, checked against current sources in October 2026. This space moves; verify a tool’s current capabilities before committing to it.

Build it in this order

Eval maturity is a sequence, not a menu. Teams that jump to the third rung before doing the first build a precise, automated measurement of the wrong thing.

  1. 00

    Vibes, plus logging

    Someone tries ten prompts by hand and forms an opinion. Legitimate during a prototype, because you do not yet know what failure looks like. The one non-negotiable is logging every input, output, tool call and retrieved chunk from day one, with a trace ID. Everything above this rung is built from those logs.

    Move up when you cannot remember whether yesterday’s prompt change made things better.

  2. 01

    Error analysis and a failure taxonomy

    Read a hundred or two real traces, label each failure in your own words, and cluster the notes into named failure modes with counts. It is the highest-value activity in the discipline and the one most often skipped, because it is manual reading and cannot be bought.

    Move up when you have named failure modes with counts, and a sense of which are worth engineering against.

  3. 02

    A golden set with programmatic assertions

    Turn the top failure modes into code. Build 100–500 cases weighted by real frequency, deliberately including the hard tail, and write deterministic checks. Run it on every change, and pin a held-out slice you never look at so you can tell overfitting from improvement.

    Move up when the failures that remain are quality judgments code cannot make.

  4. 03

    A validated LLM judge in CI

    Write a rubric of specific yes/no criteria, not “rate helpfulness 1–10.” Have people label 100–200 items, measure the judge’s agreement as Cohen’s κ, and iterate until it clears about 0.6 for most product decisions, higher when stakes are. Pin the judge version and report intervals, never bare point estimates.

    Move up when offline scores improve but you cannot tell whether users are better off.

  5. 04

    Online evaluation and monitoring

    Canary or A/B the change on live traffic with guardrails declared in advance: quality, cost per task, p95 latency, refusal and escalation rate, safety incidents. Sample production through the same judge, segmented by cohort. Feed new failures back into the golden set; a set that never changes is one you have overfit.

    Steady state. Production traces feed error analysis, which feeds the dataset, which gates the next change.

Notice what this implies about staffing. The first two rungs are careful human work on data, not model work. If an eval plan has no line item for someone reading traces, it is not a plan.

How evals lie

Each of these produces a number that is technically correct and practically misleading. They are listed roughly by how often they cost real money.

  1. 01

    The dataset nobody questions

    Sixty cases an engineer wrote from memory in the first sprint, still the release gate a year later. It measures the team’s early assumptions, faithfully, forever. Counter: Re-derive the set from sampled production traces every quarter, and keep a held-out slice you never tune against.

  2. 02

    Point estimates with no interval

    “87% vs 85%” on 200 items is almost always noise, and it gets reported as progress because nobody computed the interval. Counter: Report every number with an interval and a sample size, and pair the comparison so the interval is tight enough to mean something.

  3. 03

    The unvalidated judge

    A judge prompt written in an afternoon, never compared to human labels, producing a number quoted in planning for a year. It is consistent, which reads as reliable. Consistency is not accuracy. Counter: Measure Cohen’s κ against human labels before first use, and again whenever the judge model, rubric or product changes.

  4. 04

    Goodharting the judge

    Once a judge score becomes a target, prompts get tuned to the judge’s quirks. The score climbs; users notice nothing. Counter: Keep a human-labeled holdout the optimization loop never sees, and watch product signals for divergence.

  5. 05

    Averages that hide regressions

    One headline number stays flat while the segment your biggest customer uses falls eight points. Counter: Name the segments that matter before the test and gate on the worst one, not the mean.

  6. 06

    Benchmark scores as product evidence

    A public leaderboard decides a model choice, and the number came from a different scaffold, prompt format and retry budget than yours. Counter: Use public benchmarks to shortlist; decide on your own eval set, with every candidate run through your harness.

  7. 07

    Contamination and saturation

    Popular benchmarks leak into training data, and the frontier clusters within a point where there is no headroom left. Counter: Prefer held-out, refreshing or private sets. Treat any score in the mid-nineties as information about the benchmark, not the model.

  8. 08

    Testing the prompt, not the system

    Prompt variants get compared for weeks while the real cause is chunking, retrieval ranking, a stale index or a tool timeout. Counter: Evaluate components separately. Retrieval precision and recall, tool success rate and end-to-end quality are three different measurements.

  9. 09

    Quality-only scorecards

    A change ships on a 3-point quality gain and doubles cost per task and p95 latency. Both were measured; neither was on the gate. Counter: Put cost per resolved task, p95 latency, refusal and escalation rate, and safety incidents on the same scorecard as quality.

  10. 10

    Silent judge drift

    The judge’s underlying model is upgraded, every historical number shifts, and the trend line is quietly meaningless. Counter: Pin judge versions, re-run the baseline whenever one changes, and store the judge version with every result.

Questions to ask before you trust the number

You do not need to write a scorer to hold an eval to a standard. These separate a measurement from a vibe, and none of them require reading the code.

  1. 01

    What is in the eval set, who built it, and how do we know it looks like real traffic?

    The honest answer is often “an engineer wrote 40 cases from memory.” That is a smoke test. Ask what share came from sampled production traces.

  2. 02

    What is the confidence interval on that number?

    If nobody has one, the comparison probably cannot support the decision. The Size step above does the arithmetic.

  3. 03

    Is this judged by code, a model or a person, and if a model, what is its agreement with human labels?

    Ask for Cohen’s κ, not percent agreement, which overstates badly when one class dominates.

  4. 04

    Were both systems run on the same items?

    Paired comparisons need a fraction of the data. An unpaired comparison on a small set is usually inconclusive.

  5. 05

    Which segments got worse?

    A flat average hides a regression in the cohort that complains loudest. Ask for the breakdown before the headline.

  6. 06

    What is the p95, not the mean?

    Users experience the tail. This applies to quality as much as latency.

  7. 07

    What does each task cost, and how long does it take, at this quality?

    Quality is one axis of four: quality, cost, latency, safety. A scorecard with only quality on it is incomplete.

  8. 08

    When did the judge model or prompt last change, and was the baseline re-run?

    An unannounced judge upgrade can shift every historical number.

  9. 09

    What did the last error analysis find, and when was it done?

    If the answer is “we haven’t,” the suite is testing assumptions rather than reality.

  10. 10

    Would this suite have caught the failure that got escalated last week?

    Usually it would not, and that gap is your next sprint.

A scorecard that can carry a ship decision

The figures are illustrative. The headline would have read “quality up 2.8 points.”

MetricGraderBaselineCandidateΔ (95% CI)Gate
Output schema validcode99.4%99.8%+0.4 (−0.1, +0.9)pass
Cited record existscode96.1%98.7%+2.6 (+1.4, +3.8)pass
Answer grounded in datajudge, κ 0.7188.2%91.0%+2.8 (+0.6, +5.0)pass
Correct line item identifiedjudge, κ 0.6884.5%84.1%−0.4 (−3.1, +2.3)flat
Multi-line POs onlyjudge, κ 0.6879.0%71.2%−7.8 (−12.9, −2.7)regressed
Refuses when data missingcode + judge91.0%92.4%+1.4 (−1.0, +3.8)pass
Cost per resolved questiontelemetry$0.031$0.052+68%review
p95 latencytelemetry4.1 s7.8 s+90%breach

The rows that matter are the segment that fell 7.8 points and the latency breach. Both are invisible in an average, and both decide the launch. A scorecard’s job is to make those two rows impossible to miss.

If you do one thing

Read a hundred real traces this week and write down what went wrong, in your own words, with counts. Everything above is downstream of that list. The tools get chosen in an afternoon. The taxonomy is what takes judgment, and it is the part nobody else can do for you.

On the methods

The statistics are standard and the sources are below. The calculators implement textbook formulas and are exact for their stated assumptions. The judge lab is an illustrative simulation: it shows the direction and rough size of each effect, not a measurement of any model. Everything runs in your browser; nothing you enter is sent anywhere.

Sources