# EvalShift blog — full text for AI tools > Every post from https://www.evalshift.dev/blog, complete, in publication order (newest first). > The web pages at those URLs render this same markdown. --- ## This one passed: 71% cheaper, and the agent still does the same thing URL: https://www.evalshift.dev/blog/what-a-passing-migration-proves Published: 2026-09-16 Tag: migration Summary: An EvalShift migration report that said PASS: a 71% cheaper model inside every policy budget. Where each threshold came from, and what PASS does not claim. Takeaways: - EvalShift replayed 120 captured examples against the source and candidate model and returned PASS: 71.2% cheaper, 61.9% faster, all ten policy budgets held. - EvalShift's migration_policy in evalshift.yaml holds seven budgets, and a slices map tightens any of them per slice — here security got zero regressions and zero tool divergence, refund got zero argument drift. - Check every rate budget against the suite size first: a rate over 108 rows moves in steps of 0.93%, so a 1% budget is zero tolerance in disguise, and EvalShift warns when a budget is sub-granular. - EvalShift judges proportion budgets with a 95% Wilson interval, asymmetrically: a held budget is conclusive however wide the interval, and only a breach needs the interval to confirm it. - PASS does not mean nothing changed: 16 comparisons regressed and 28 improved, and EvalShift's per-example diff shows each one. It means every change stayed inside limits fixed before anyone saw a number. The previous post was about a migration that was cheaper, faster and still failed. This is the other outcome. It is the more common one once a suite is in decent shape, and I see it written up far less often, because a pass is boring. It shouldn't be. A pass is only worth something if the limits were fixed before the run. So this post is the whole thing: the policy, the reasoning behind every number in it, and the EvalShift report it produced. For anyone arriving here cold: EvalShift is a local-first CLI for testing LLM model migrations. It replays a frozen golden suite against your current model and a candidate, scores every pair of outputs with tool-call, structural, semantic and LLM-judge evaluators, runs paired statistics over the deltas, and turns a `migration_policy` block you wrote in `evalshift.yaml` into one of four verdicts: `pass`, `conditional_pass`, `fail` or `inconclusive`. It writes a single-file `report.html` on your machine; nothing is uploaded unless you run `evalshift push`. Every figure in this post is a panel from that report. The candidate: a customer-support agent with six tools (`lookup_customer`, `lookup_order`, `check_refund_policy`, `issue_refund`, `escalate_to_human`, `search_kb`), moving from `gemini-3.7-pro` to `gemini-3.7-flash`. The suite: 120 examples recorded in production by the EvalShift capture SDK, promoted into a golden JSONL suite with `evalshift capture sync`, and sliced by tag according to what the conversation was about. | Slice | Examples | What is in it | | --- | --- | --- | | `routine` | 42 | order status, shipping, account questions | | `refund` | 26 | refund and return requests | | `security` | 24 | account access, password and payment-method changes | | `customer_lookup` | 16 | requests that need a customer record first | | `text_only` | 12 | greetings, thanks, off-topic | One `evalshift compare` command, real API calls on both sides, and the report opened on: ```report { "kind": "verdict", "verdict": "pass", "summary": "10 of 10 budgets within policy.", "rates": { "equivalent": 90.1, "improved": 6.3, "regressed": 3.6 }, "cards": [ { "eyebrow": "Advisory signal", "value": "3.6%", "unit": "regression rate", "tone": "ok", "note": "Below the max_overall_regression_rate of 5%. 16 of 444 scored comparisons, none above medium severity." }, { "eyebrow": "Economics", "value": "-71.2%", "unit": "cost", "tone": "ok", "note": "$1.9248 → $0.5544. Latency -61.9%. Both inside +0% cost / +30% latency." } ], "caption": "EvalShift's verdict card. Ten budgets, all held; the regression rate and the economics sit beside it so nobody has to scroll to find out what the pass cost." } ``` ```report { "kind": "strip", "cells": [ { "label": "Examples", "value": "120" }, { "label": "Calls", "value": "240", "note": "0 cached" }, { "label": "Failed / truncated", "value": "0 / 0", "tone": "ok" }, { "label": "Total cost", "value": "$2.4792" }, { "label": "Latency Δ", "value": "-61.9%", "tone": "ok" }, { "label": "Avg score Δ", "value": "+0.011", "tone": "ok" } ], "caption": "The run strip. Same 120 examples on both sides, nothing failed or truncated, so every comparison below is over the full suite. EvalShift excludes truncated and errored calls from the statistics, so this row is worth checking first." } ``` Cheaper by 71.2%, faster by 61.9%, verdict PASS. The economics look like the last post's. The difference is everything that was decided before the run. ## The numbers were written before the run EvalShift's `migration_policy` is a block in `evalshift.yaml` with seven budgets: overall regression rate, critical regression count, equivalent-or-better rate, tool-argument drift, tool-selection divergence, cost increase and latency increase. A `slices` map under it lets any slice override any budget, inheriting the top-level value where it doesn't. Evaluators are configured per suite; budgets are set once and tightened per slice. `evalshift init --profile cost-reduction` scaffolds a starting policy: 2% overall regression rate, zero critical regressions, 97% equivalence, 1% tool-argument drift, 5% cost increase, 30% latency increase. That is a starting point, not a decision. This is what the project actually ran with: ```yaml migration_policy: max_overall_regression_rate: 0.05 max_critical_regressions: 0 min_equivalence_rate: 0.90 max_tool_argument_drift: 0.10 max_tool_divergence: 0.05 max_cost_increase: 0.0 max_latency_increase: 0.30 slices: security: max_overall_regression_rate: 0.0 max_tool_divergence: 0.0 refund: max_tool_argument_drift: 0.0 ``` Three things moved between the profile and this file. ### Check every rate against the suite size A rate over *n* rows can only move in steps of 1/*n*. This suite has 108 tool-argument rows, so one drifted row is 0.93%, and a 1% budget is a zero-tolerance budget wearing a percentage. EvalShift prints a recommendation when a budget is below the granularity of its denominator, naming the budget, the value and the row count, and this one would have triggered it. So the rule I use: if I mean zero, I write `0.0`. Where I don't, I set a number the sample can resolve. 10% argument drift on 108 rows is ten reworded search queries, which is what drift on this agent looks like in practice. ### Tolerance where wording lives, zero where money and access live The overall regression rate went *up*, from 2% to 5%. The pairwise LLM judge is blocking on this project, and a judge flips on rewording; 2% of 444 comparisons is nine judge calls deciding the migration. 5% leaves room for the judge to disagree about phrasing without anything real being allowed through. The strictness moved into the slices instead. `security` gets zero regressions and zero tool-selection divergence: a model that starts routing an account-access request to a different tool does not get a percentage. `refund` gets zero argument drift: an order id or an amount that drifts is a wrong refund, not a reworded one. The slices that are *not* in the policy matter too. `customer_lookup` has 16 examples. EvalShift tests a comparison below 20 pairs but flags it uncertain, so that slice inherits the top-level budgets and gets no tighter ones. It is the slice I would grow before I tightened it. ### The cost budget is the reason for the migration `max_cost_increase: 0.0`. The point of the exercise is to spend less. A candidate that costs more has failed before any quality number is read, and a policy should say so instead of leaving it to whoever reads the economics card. EvalShift measures cost and latency from the run's own calls, so these two budgets gate even when no quality evaluator does. Two evaluator decisions sit alongside the budgets. Every EvalShift evaluator carries a `blocking` flag; advisory (`blocking: false`) results are reported and ranked but never change the verdict. The judge is `blocking: true` here; `init` writes `false` because at a dozen examples judge noise would decide the verdict, and at 120 with an audited criterion it earns its vote. `semantic` stays advisory. It measures wording, and wording is the one thing this migration was allowed to change. ## What the run measured ```report { "kind": "budgets", "rows": [ { "name": "Overall regression rate", "id": "max_overall_regression_rate", "scope": "overall", "observed": "3.6%", "limit": "≤ 5.0%", "tone": "ok", "note": "16 of 444 · 95% CI 2.2–5.8%" }, { "name": "Critical regressions", "id": "max_critical_regressions", "scope": "overall", "observed": "0", "limit": "≤ 0", "tone": "ok", "note": "of 444" }, { "name": "Equivalent-or-better rate", "id": "min_equivalence_rate", "scope": "overall", "observed": "96.4%", "limit": "≥ 90.0%", "tone": "ok", "note": "428 of 444" }, { "name": "Tool-argument drift", "id": "max_tool_argument_drift", "scope": "overall", "observed": "4.6%", "limit": "≤ 10.0%", "tone": "ok", "note": "5 of 108 tool-argument rows" }, { "name": "Tool-selection divergence", "id": "max_tool_divergence", "scope": "overall", "observed": "2.8%", "limit": "≤ 5.0%", "tone": "ok", "note": "3 of 108 divergence rows" }, { "name": "Cost increase", "id": "max_cost_increase", "scope": "overall", "observed": "0.0%", "limit": "≤ 0.0%", "tone": "ok", "note": "cost fell 71.2%" }, { "name": "Latency increase", "id": "max_latency_increase", "scope": "overall", "observed": "0.0%", "limit": "≤ 30.0%", "tone": "ok", "note": "latency fell 61.9%" }, { "name": "Overall regression rate", "id": "max_overall_regression_rate", "scope": "security", "observed": "0.0%", "limit": "≤ 0.0%", "tone": "ok", "note": "0 of 96" }, { "name": "Tool-selection divergence", "id": "max_tool_divergence", "scope": "security", "observed": "0.0%", "limit": "≤ 0.0%", "tone": "ok", "note": "0 of 24" }, { "name": "Tool-argument drift", "id": "max_tool_argument_drift", "scope": "refund", "observed": "0.0%", "limit": "≤ 0.0%", "tone": "ok", "note": "0 of 26" } ], "caption": "Every budget against its limit, as EvalShift reports them. Seven at the top level, three on the two slices where a regression is a wrong action rather than a reworded one." } ``` The first row is the one worth reading twice. 16 of 444 is 3.6%, and its 95% Wilson interval runs from 2.2% to 5.8%. The interval crosses the limit. EvalShift's rule is asymmetric on purpose. A breach fails only when the interval confirms it; a breach the interval cannot confirm returns `inconclusive`, because the suite was too small to say. A budget the observation held is conclusive however wide its interval, because a wide interval must never downgrade a clean run. This budget held, so it passes, and the interval is printed so the reader knows how much room there was. The three slice budgets held at exactly zero. Zero on 96 comparisons means something. Zero on 12 would not, which is why `text_only` has no slice budget at all. ## By evaluator Four evaluators scored this run. EvalShift's `tool_selection` evaluator reads the recorded traces and scores two axes: conformance, where each side is graded against the suite's recorded tool calls, and divergence, where the target is graded against what the source did. `tool_arguments` scores argument values field by field against the expected call. The `llm_judge` is pairwise and sees the two outputs as anonymous A and B. `semantic` is embedding similarity between the two outputs. The first three are blocking; the last is advisory. ```report { "kind": "evaluators", "rows": [ { "name": "Routing — conformance", "id": "routing · tool_selection.conformance", "axis": "each side graded against the suite's recorded tool calls", "n": 108, "delta": "+0.046", "effect": "0.28", "magnitude": "small", "ci": "[0.09, 0.47]", "confidence": "likely", "severity": "improved" }, { "name": "Routing — divergence", "id": "routing · tool_selection.divergence", "axis": "the target graded against what the source did", "n": 108, "delta": "-0.028", "effect": "0.17", "magnitude": "negligible", "ci": "[-0.36, 0.02]", "confidence": "unclear", "severity": "none" }, { "name": "Routing args", "id": "routing_args", "n": 108, "delta": "-0.004", "effect": "0.03", "magnitude": "negligible", "ci": "[-0.22, 0.16]", "confidence": "unclear", "severity": "none" }, { "name": "LLM judge: equivalence", "id": "llm_judge.equivalence", "n": 120, "delta": "+0.029", "effect": "0.12", "magnitude": "negligible", "ci": "[-0.06, 0.30]", "confidence": "unclear", "severity": "none" }, { "name": "Semantic similarity", "id": "semantic", "advisory": true, "n": 120, "delta": "-0.041", "effect": "0.44", "magnitude": "small", "ci": "[-0.62, -0.26]", "confidence": "likely", "severity": "medium", "blurb": "Reported, not gating: blocking is false." } ], "caption": "Overall, by evaluator. Each row is a paired test over that evaluator's deltas: effect size with a 95% interval, and a confidence label from the Benjamini-Hochberg corrected p-value. Four blocking rows say equivalent or improved; the one regression is on the advisory evaluator that measures wording." } ``` Four blocking rows say equivalent or improved. The one that regressed is advisory, and it measures the one thing this migration was allowed to change. Had `semantic` been blocking, the same run would have come back `conditional_pass` on a medium-severity regression in phrasing. That is not a reason to make every evaluator advisory. It is a reason to decide, per evaluator, whether what it measures is something you are willing to block a migration on. The conformance row says the candidate matched the suite's recorded tool calls *more often* than the model that produced them: 11 examples improved, one regressed. Most of the eleven were refund requests where the source went straight to `issue_refund`. ## The diffs I still read A pass is not permission to skip the diff. For every flagged example, the EvalShift report shows the reason it was flagged, the tool calls on each side, the argument-level diff, and the conversation context that led to it. Two examples from this run, one from each side of the ledger. ```report { "kind": "example", "id": "cap_6b1e40f2a9c34d0b8e7d2a5f31c9e804", "turn": 1, "what": "Routing — conformance", "delta": "+0.500", "tone": "ok", "why": { "scores": "source 0.50 → target 1.00 (0–1)", "label": "IMPROVED", "text": "The suite expected check_refund_policy before issue_refund. The source went straight to issue_refund; the target called both, in order." }, "tools": { "source": "issue_refund", "target": "check_refund_policy, issue_refund" }, "source": { "score": "0.500", "tone": "bad", "calls": [ { "tool": "issue_refund", "args": "{\"order_id\": \"ord_58213\", \"amount\": 42.0, \"reason\": \"damaged on arrival\"}" } ], "final": "Done — I've refunded $42.00 to your original payment method. You'll see it in 3–5 business days." }, "target": { "score": "1.000", "tone": "ok", "calls": [ { "tool": "check_refund_policy", "args": "{\"order_id\": \"ord_58213\"}", "mark": "extra" }, { "tool": "issue_refund", "args": "{\"order_id\": \"ord_58213\", \"amount\": 42.0, \"reason\": \"damaged on arrival\"}" } ], "final": "Your order is inside the 30-day window, so I've refunded $42.00 to your original payment method. Expect it in 3–5 business days." }, "diffs": [ "extra tool: Target added check_refund_policy at position 1.", "matched: issue_refund arguments are identical on both sides." ], "caption": "An improvement, as EvalShift's per-example diff shows it. The recorded ground truth expected the policy check first; the source model that produced the traffic skipped it and the candidate did not." } ``` The refund went out either way, and the final text on both sides would pass any output check you care to write. The trace is where the difference lives: one side verified before acting and the other did not. This is the class of change a text evaluator cannot see in either direction, and the reason EvalShift's tool-call evaluators score the trace rather than the prose. ```report { "kind": "example", "id": "cap_9d27c5e1b04a4f6e9a3c7d8e2f105b6c", "turn": 0, "what": "Routing args", "delta": "-0.160", "tone": "bad", "why": { "scores": "source 1.00 → target 0.84 (0–1)", "label": "ARGUMENT_VALUE_DRIFT", "text": "search_kb.query scored 0.84 against the expected value under the semantic strategy, below the 0.9 drift floor. Same tool, same intent, different wording." }, "tools": { "source": "search_kb", "target": "search_kb" }, "source": { "score": "1.000", "tone": "ok", "calls": [ { "tool": "search_kb", "args": "{\"query\": \"change shipping address after order placed\"}" } ], "final": "You can change the address until the order ships — open the order and tap “Edit address”. Once it's marked shipped, reply here and I'll reroute it with the carrier." }, "target": { "score": "0.840", "tone": "bad", "calls": [ { "tool": "search_kb", "args": "{\"query\": \"edit delivery address existing order\"}" } ], "final": "Until the order ships you can edit the address yourself from the order page. If it has already shipped, let me know and I'll set up a carrier reroute." }, "diffs": [ "argument drift: search_kb.query scored 0.84 against the expected query (drift floor 0.9).", "same tool set: no calls added or removed." ], "caption": "A regression the budget was built to tolerate. Five of 108 tool-argument rows drifted like this one; the limit was ten." } ``` A reviewer would call those the same query. The evaluator scored it 0.84 against a 0.9 floor and counted it as drift, five rows out of 108 did the same, and the budget said five is fine. That is what the 10% was for. The same five rows in the `refund` slice would have failed the run, and that was also decided in advance. ## What EvalShift did in this run The whole run, as a list of the product's parts, in the order they were used: - **The capture SDK** recorded the agent's real conversations in production, tool calls included, and `evalshift capture sync` promoted them into a frozen golden JSONL suite with the tool evaluators written from what the captures actually contained. - **`evalshift compare`** replayed every example against the source and the target model, paired per example, with the same inputs, tools and context on both sides. - **The tool-call evaluators** (`tool_selection`, `tool_arguments`) scored the traces, the **pairwise LLM judge** scored the outputs, and **`semantic`** measured drift in wording, advisory only. - **Paired statistics** turned each evaluator's deltas into an effect size, a 95% interval and a corrected confidence label, so a two-point average drop and a real regression are told apart mechanically. - **`migration_policy`** in `evalshift.yaml` held seven budgets, three of them tightened to zero on the `security` and `refund` slices, and every proportion budget was judged with a Wilson interval that can return `inconclusive` instead of a false fail. - **`report.html`** was written locally, verdict first, with a per-example diff for everything flagged. Nothing left the machine. - **`--policy-gate`** made the verdict an exit code, which is what the EvalShift GitHub Action uses to block a pull request on the same policy once the suite runs in CI. ## What a pass proves Not that nothing changed. 16 comparisons regressed and 28 improved; the agent phrases search queries differently and checks the refund policy more often than it used to. It proves that every change stayed inside limits that were written down before anyone saw a number, on a suite that was frozen before the run. That is the entire claim, and it is enough to act on, because there is nothing left to negotiate: the argument about what counts as acceptable happened in the YAML, not in the meeting after the report. What happens next is the boring part, which is the point. The model string changes in production. The suite does not. It runs again on the next pull request through the EvalShift GitHub Action, gated on the same policy, against the new baseline. ```bash evalshift compare --suite-name support_routing --to gemini-3.7-flash --policy-gate --open ``` If the previous post was the reason to run the comparison, this one is the reason to write the policy first. --- ## The new model was 59% cheaper and 75% faster. I still wouldn't ship it. URL: https://www.evalshift.dev/blog/cheaper-faster-and-still-a-fail Published: 2026-09-12 Tag: migration Summary: A candidate model cut cost 59% and latency 75% with zero failed calls. The migration report still said FAIL, because the agent's behavior changed. Takeaways: - Cost, latency and error rate all improved — 58.8% cheaper, 75.2% faster, zero failed or truncated calls — and the migration report still returned FAIL. - The overall regression-rate budget passed. Tool-selection divergence at 25% against a 10% limit and an equivalence rate of 61.4% against a 75% floor did not. - The regressions were behavioral: an extra search_web call instead of an answer, an action reported as done with no tool call, a request routed to a different tool. - Per-evaluator results disagreed — routing arguments and conformance held while tool-selection divergence, semantic similarity and the pairwise judge regressed. One averaged score would have hidden that. - Economics alone cannot tell you whether a migration is safe. Replay the same suite against both models and read the per-example diff before the model string changes. The migration looked like an obvious win on economics. In one EvalShift run, the candidate model was: - 58.8% cheaper - 75.2% faster - 0 failed or truncated calls ![Run economics table: source and target each made 16 calls with no failures or truncations; the target cost $0.0741 against $0.1799 and averaged 773 ms against 3.1 s.](/blog-images/blog1/blog1-3.png "Run economics. Same 16 calls on both sides, no failures, and every number a finance reviewer cares about moved the right way.") Then the migration report returned: **FAIL** Not because the model crashed. Not because the API changed. Not because the output stopped parsing. It changed what the agent actually did. ![Migration verdict card: FAIL, 5 of 7 budgets within policy, 61.4% equivalent, 25.0% improved, 13.6% regressed, cost −58.8%, latency −75.2%.](/blog-images/blog1/blog1-1.png "The verdict card. Five of seven budgets passed; the two that did not are the two that describe behavior.") ## What failed I replayed 16 real examples against the source and candidate models. Across the run, 13.6% of evaluated comparisons regressed. The overall regression-rate budget actually passed. Two other checks did not: - Tool-selection divergence hit 25%, against a 10% limit. - The overall equivalence rate fell to 61.4%, below the required 75%. ![Report verdict and findings: the target breached max_tool_divergence at 25.0% against a 10% ceiling and fell below min_equivalence_rate at 61.4% against 75%; findings list searching instead of answering, confirming actions it never performed, and routing to the wrong tool.](/blog-images/blog1/blog1-2.png "The written verdict and its findings. The recommendation is not to migrate until the tool-selection and semantic regressions are fixed.") That distinction matters. Looking only at latency, cost, successful requests, or even a few manually inspected outputs would have made this migration look pretty attractive. The behavioral diff told a different story. ## The failures weren't cosmetic One example asked for an opinion on working late. The source model answered the question directly. The candidate instead issued a `search_web` call and returned no text. ![Per-example diff for the working-late question: the source called no tools and answered in two sentences; the target called search_web with the query "is working late worth it productivity well-being" and produced no final text. Flagged as TOOL_SELECTION_DRIFT.](/blog-images/blog1/blog1-5.png "The diff for that example. Source trace on the left, target trace on the right, and the reason it was flagged at the top.") Other cases had the same general problem in different forms: - an action was reported as completed even though the corresponding tool was never called; - a request was routed to a different tool; - an unnecessary tool call appeared where the source answered directly. These are not necessarily signs that the candidate model is "bad." They show that changing the model changed the application. For an agent, the model is part of the control flow. ## Aggregate quality can hide this The evaluator breakdown made the tradeoff clearer. Routing arguments and routing conformance were broadly equivalent in this run, while tool-selection divergence, semantic similarity, and the pairwise equivalence judge showed regressions. ![Per-evaluator table over 16 examples: routing divergence regressed (score delta −0.167, likely), routing args equivalent (+0.179, unclear), routing conformance equivalent (+0.047, unclear), the LLM equivalence judge regressed critically (−0.438, certain), and semantic similarity regressed (−0.214, likely).](/blog-images/blog1/blog1-4.png "Overall, by evaluator. Two rows say equivalent, three say regressed, and each comes with its own effect size, confidence interval and confidence label.") A single average score would flatten all of that into one number. For a migration, I care more about the question: *What changed, on which examples, and is that change acceptable for this application?* ## This is why I built EvalShift I kept running into model migrations that were tested roughly like this: change model → try several prompts → outputs look fine → ship. That works until the difference is something subtle like an extra tool call, a missing action, or different routing behavior. EvalShift replays the same suite against the source and target model, compares the behavior, and produces the migration decision and individual diffs before the model string gets changed in production. The CLI is local-first and produces a self-contained HTML report; hosted upload is optional. The interesting result from this run wasn't that the candidate was worse. It was that "59% cheaper and 75% faster" wasn't enough information to decide whether the migration was safe. --- ## What actually breaks when you switch LLMs URL: https://www.evalshift.dev/blog/what-breaks-when-you-switch-llms Published: 2026-08-20 Tag: migration Summary: A model swap is a behavior change with no diff to review. What moves besides the answer text, and why the worst regressions read as correct. Takeaways: - A model swap changes behavior — tools called, call order, arguments, output shape, refusals, latency, cost — and ships through a config file with no diff for anyone to review. - The expensive regressions are invisible in the output: an agent that skips issue_refund and still says the refund was processed passes every text-based check. - Migrations usually edit the prompt too, which puts two variables in one comparison — run new model with the old prompt first, then tune, so each result means something. - Hand-written eval cases cover the interactions you already handle; the failures live in recorded traffic you would never have thought to write down. - Significance is not importance: a 0.02 semantic delta can be real and irrelevant, one missing verify_payment call can be statistically invisible and serious. Write the policy before you see the numbers. Changing the model behind an AI feature looks like the smallest change you will make all week. One string moves from `gpt-x` to `gemini-y`. The new model is cheaper, or faster, or ahead on the benchmark someone linked in Slack. You try a handful of prompts, the answers read fine, you ship. The reason this keeps going wrong is that a model swap is not a config change. It is a behavior change, delivered through a config file, with no diff for anyone to review. ## The change surface is bigger than the answer text Swapping the model can move any of these independently, and most teams only look at the last one: | What moves | How it usually surfaces | | --- | --- | | which tools the agent calls | a step silently stops happening | | the order of those calls | a check runs after the action it was meant to gate | | tool arguments | right tool, wrong amount, wrong id, wrong units | | structured output | your parser throws, or worse, doesn't | | refusal behavior | the model declines work it used to do | | verbosity and format | downstream regex and UI truncation start missing | | latency | p95 doubles, nobody attributes it to the swap | | tokens and cost | the cheaper model turns out to be the pricier one per task | | answer quality | the only one the playground actually shows you | A migration can improve one row and wreck another. A model that answers just as well but issues one extra tool call per turn is not a cost reduction. You will not learn that by reading answers. ## The worst regressions are invisible in the output Take an agent that is supposed to do this: ```text lookup_order("A-339") issue_refund("A-339", 29.99) ``` The new model does this instead: ```text lookup_order("A-339") ``` and replies: > Your refund has been processed. Every text-based check passes. The sentence is fluent, on topic, and exactly what the old model said. The refund did not happen. If you are scoring outputs, this regression is not merely hard to catch — it is invisible by construction, because the output is correct and the behavior is not. The same shape covers most of the expensive failures: the `verify_payment` call that stops firing, the retry loop that starts, the confirmation step that moves after the write. What changed is the trace. The prose stayed still. ## You are usually changing two things at once Prompts are coupled to models. The system prompt in production has been tuned — often over months, often by accretion — against one model's quirks. Point it at a different model and some of that tuning becomes dead weight and some becomes actively harmful. So the honest migration usually involves editing the prompt too, and now the comparison has two independent variables in it. When quality moves, nobody can say whether the model did it or the rewrite did. The fix is boring: change one thing per run. Baseline old model with old prompt. Run new model with old prompt — that is the model's effect, unflattering as it may be. Then tune the prompt for the new model and run again, against the same frozen cases. Two comparisons, each interpretable, instead of one that isn't. ## Hand-written test cases test the paths you already handle Writing eval prompts by hand feels productive and produces a suite shaped like your mental model of the product. That is the problem. The prompts you can think of are the interactions you already understand well enough to have handled. Real failures come from the inputs you would never have written down: the eleven-turn conversation where a constraint set in turn three quietly expires, the message that arrives with half the context missing, the turn right after a tool returned an error, the customer typing in a language your template never anticipated. You cannot reconstruct those from memory. You have to record them. Which gives a workflow, independent of what you use to run it: ```text real agent behavior ↓ capture representative cases ↓ freeze a golden suite ↓ baseline model vs candidate model, same inputs ↓ compare behavior, not just text ↓ ship or reject ``` The freezing step is the one people skip. If cases are still being edited while models are being compared, two things are moving and the diff between them describes neither. ## An average is not a verdict Model A scores 0.91, model B scores 0.89, and someone screenshots it into the migration thread. That number cannot carry the decision. With a couple of dozen noisy samples, a two-point gap is well inside what you would see running the *same* model twice. Two things make it a real comparison. Run paired — every case against both models, then subtract per case, so the fact that some cases are inherently harder cancels instead of drowning the signal. And correct for multiple comparisons — a suite producing forty comparisons will hand you two significant findings by luck alone, so something like Benjamini-Hochberg has to sit between the tests and the conclusion. But statistical significance is not product importance, and this is where the reasoning usually stops one step early. A semantic-similarity drop of 0.02 can be real, reproducible, significant, and irrelevant. One missing `verify_payment` call across two hundred cases is statistically nothing and operationally a serious problem. Significance tells you the effect exists. It has no opinion about whether you should care. ## Decide what matters before you see the numbers Which is why the last artifact of a migration is a written policy — thresholds agreed while the result is still unknown, so the verdict is read off rather than negotiated: ```text missing critical tool -> fail invalid JSON above threshold -> fail latency +10% -> warn small semantic delta -> ignore ``` Write that after the run and the thresholds bend around the number you were hoping for. Everyone does this; nobody means to. A policy also gives you a fourth verdict worth having explicitly: *inconclusive*. Not "pass", not "fail" — "this suite is too small to tell you". Teams that collapse that into a pass ship regressions they had the evidence to catch, one underpowered comparison at a time. ## Disclosure, and the part that survives it I build [EvalShift](/), which does exactly this: capture real agent runs, freeze them into a golden suite, run both models paired, score tool calls and arguments and structure alongside output quality, and gate the pull request on a policy you wrote in advance. So take the tool mention as interested. The argument underneath it is not: If an LLM decides behavior in your system, then changing the model is a behavior change, and it deserves the same treatment as a dependency bump that alters runtime semantics — frozen test cases, a before-and-after run, and a decision rule written while you still have no stake in the answer. Not a string edit in a config file, followed by hope. ## Keep reading - [How to test an LLM model migration before you ship it](/blog/test-llm-model-migration-before-you-ship) — the same argument as a method, step by step. - [Evaluating agent tool calls](/blog/evaluating-agent-tool-calls) — what text evaluators structurally cannot see, and the four evaluators that can. - [Build a golden eval suite from production traffic](/blog/build-a-golden-suite-from-production-traffic) — turning recorded runs into the frozen suite this all depends on. --- ## Prompt edits deserve the same gate as a model swap URL: https://www.evalshift.dev/blog/prompt-regression-testing Published: 2026-08-13 Tag: prompts Summary: A prompt edit changes behavior the way a model swap does. Where the prompt text lives decides whether CI diffs it or passes on silence. Takeaways: - A prompt edit is a behavior change with no diff to review — the same golden suite that gates a model swap gates it, unchanged. - eval_config_hash covers version, prompts, defaults, evaluators and slices, and the baseline lookup skips every run whose hash differs. - detection: manual puts the prompt body inside that hash, so every wording edit re-baselines the project and the PR check passes with no comparison at all. - detection: python_string keeps the body in your application code — the hash holds still, the diff runs, and a regression reads as a regression. - A permanently green EvalShift check usually means no comparable baseline was found, not that the suite passed. A model swap gets a meeting. A prompt edit gets a commit. Both change what your product says to users, and only one of them is reviewed like it matters — usually because a prompt diff is one paragraph of English that reads fine and proves nothing. The gate you already built for model migrations works on prompt edits unchanged: a golden suite, a paired run, a diff against the base branch. But whether a prompt edit actually reaches that gate depends on something most teams never look at — where the prompt text lives. ## The gate, briefly On a pull request the EvalShift action runs the suite, uploads the run, asks the Cloud API for a compatible baseline run on the base branch, fetches the server-side diff, updates one PR comment, and sets the `evalshift/regression` commit status. Under `fail-on: regression` the check goes red when the diff shows regressions. The load-bearing word is *compatible*. The lookup scans recent `available` runs on the base branch for the same suite, and skips every one whose `eval_config_hash` differs from the candidate's. No compatible baseline means no diff — and no diff means the check passes. ## Where the prompt lives decides whether it is gated `eval_config_hash` is a SHA-256 over the canonical JSON of your config snapshot: `version`, `prompts`, `defaults`, `evaluators`, `slices`. The `prompts` block is in the hash. What is inside that block depends on the detection mode: ```yaml # detection: manual — the prompt body IS the config prompts: - id: replay detection: manual content: "You are a support agent. Answer in at most three sentences.\n\n{input}" variables: [input] # detection: python_string — the config only points at the body prompts: - id: customer_routing detection: python_string path: prompts.py variable: AGENT_SYSTEM_PROMPT variables: [query] ``` With `detection: manual`, editing one word of the prompt changes `prompts[].content`, which changes the config snapshot, which changes `eval_config_hash`. Your PR's run is now incomparable with every run on `main`. The action finds no compatible baseline, reports exactly that in the comment, and exits green. With `detection: python_string`, `content` must be null — the config carries only `path` and `variable`, and the body is AST-extracted from the Python file at run time. Edit the prompt and the hash does not move: the baseline lookup matches, the server diffs the runs directly, and the regression shows up as a regression. > A permanently green EvalShift check usually means no comparable baseline exists, not that the > suite is passing. If the check went green the same week someone rewrote the system prompt inline, > that is the mechanism, not a coincidence. The same rule governs everything else in the hash: retuning `defaults`, adding an evaluator, or renaming a slice all re-baseline you on purpose. Prompt wording is the one that gets edited weekly and the one nobody expects to sever the comparison. ## What still changes when the text changes Keeping the prompt out of the config hash does not make edits invisible to the pipeline: - The response cache is keyed on `{model, prompt, inputs, temperature, max_tokens[, history]}`, so an edited prompt is a cache miss. You cannot accidentally score new wording with old outputs. - Every example is validated against every prompt — template variables covered — *before* any model call is dispatched. Dropping `{order_id}` from the template fails the run in seconds rather than after $9 of calls. - `python_string` extraction never imports or executes your code. It AST-walks for a module-level string assignment and takes the last one; f-strings, concatenation, `.format()`, and function calls are rejected rather than evaluated. That last rule has a sharp edge. A prompt assembled at runtime cannot be extracted, and the documented workaround is to paste it into the config as `detection: manual` — which puts the body back inside the hash and back into re-baselining on every edit. Prefer refactoring the template so the literal is the literal and the variables are `{placeholders}` the suite fills in. ## Wire it once Keep one copy of the prompt, in the application code that ships it, and point the config at it: ```yaml prompts: - id: customer_routing detection: python_string path: app/prompts.py variable: AGENT_SYSTEM_PROMPT variables: [query] ``` Your app imports `AGENT_SYSTEM_PROMPT`; EvalShift reads the same file. There is no second copy to drift, and a reviewer looking at the PR diff sees the prompt change and the eval result side by side. Tool schemas follow the same no-second-copy rule by a different route: they live on the suite examples, not in the config — `evalshift capture sync` records each example's toolset as a content-addressed `toolset_ref` sidecar — so they never enter `eval_config_hash` either. Then let the run answer the question the prose cannot. Both sides score paired per example, so "the new prompt is more concise" becomes a delta per evaluator with a severity attached, and the migration policy turns it into `pass`, `conditional_pass`, `fail`, or `inconclusive` without a meeting. Roll out with `fail-on: never` for a week to collect baselines, then switch to `fail-on: regression` once the comments match your judgement. ## The review question worth adopting For every PR that touches a prompt: *did this run diff against a baseline, or against nothing?* The comment answers it in one line. A prompt edit that produced no comparison has not been tested — it has only been observed to compile. ## Keep reading - [LLM regression testing in CI](/blog/llm-regression-testing-in-ci) — the gate itself, end to end. - [How many eval cases do you need?](/blog/how-many-eval-cases-do-you-need) — sizing the suite that gate reads. - [Evaluating agent tool calls](/blog/evaluating-agent-tool-calls) — what changes in the trace when a system prompt changes. - [Configuration](/docs/configuration) and [Baselines](/docs/baselines) — the config fields and the baseline model in full. --- ## How many eval cases do you need? URL: https://www.evalshift.dev/blog/how-many-eval-cases-do-you-need Published: 2026-08-12 Tag: statistics Summary: Suite size is the wrong question: n is counted per prompt, evaluator and slice. What the analysis drops, and how many pairs a verdict needs. Takeaways: - n is counted per (prompt_id, evaluator_name, slice_name) comparison, not per suite — 200 examples across 3 prompts, 4 evaluators and 3 slices is up to 36 separate tests. - Below 5 paired observations a comparison is skipped as insufficient; between 5 and 20 it is tested but flagged uncertain. Target 20 per comparison you would act on. - Rows vanish before n is computed: non-applicable rows, max_tokens truncations, and errored evaluator calls all leave the sample. - Slicing splits the same rows into smaller groups and raises the Benjamini-Hochberg correction burden across the run — slice only where you would act on the subgroup. - Rate budgets are Wilson-CI aware, so a small suite returns inconclusive rather than failing: a 3% regression budget cannot be confirmed breached by 30 examples. Everyone asks the question the same way: how many examples does a golden suite need? Fifty? Two hundred? The number that actually decides whether your run can say anything is not the size of the suite. It is `n` inside one comparison — and the suite is chopped into comparisons before a single statistic is computed. This is where most first suites disappoint. Two hundred examples feel serious, come back `inconclusive`, and nobody can explain why. Here is how `n` is counted, what is dropped before counting, and how to size a suite so the verdict is earned rather than lucky. ## `n` is per comparison, not per suite The analysis groups paired deltas — `target_score - source_score`, per example — by the triple (`prompt_id`, `evaluator_name`, `slice_name`). Every one of those groups is tested on its own, and every one has its own `n`. So a 200-example suite run across 3 prompts, scored by 4 evaluators, sliced 3 ways is not one comparison with n=200. It is up to 36 comparisons, each drawing on the rows that match its slice. Give the smallest slice 15 examples and that column of the matrix is stuck at n=15 no matter how impressive the suite total looks. Three constants decide what happens next: - `MIN_N_FOR_TEST = 5` — below five paired observations the comparison is skipped entirely, with severity `insufficient`. It is not scored as "no change"; it is scored as "not measured". - `MIN_N_RELIABLE = 20` — between 5 and 20 the comparison is tested but flagged uncertain. - Zero variance (std < 1e-9) — skipped with severity `none`, because every delta was identical. Twenty per group is the line where a result stops carrying an asterisk. Work backwards from the groups you care about, not forwards from a round suite size. ## Rows evaporate before `n` is counted The `n` a comparison is tested at is smaller than the number of examples you wrote, and the gap is deliberate: - Rows an evaluator measured nothing on — a tool-only turn handed to a text evaluator, say — write no score at all, so they are absent from `n` before it is computed, and noted as "K of N rows not applicable". They also leave the slice aggregates and the policy metrics. - Truncated calls — the ones that hit `max_tokens` — are excluded from statistics, because a cut-off answer scores badly for a reason that has nothing to do with the model's quality. - Evaluator-side failures (a judge call that broke, an embedding request that errored) are stored as `errored` and excluded. A flaky judge shrinks `n` and drifts a comparison toward `insufficient`, which is visible, instead of poisoning the mean, which is not. > When nothing survives, the comparison reports n=0 with a note prefixed `nothing measured:` — > never severity `none`. Unmeasured is not equivalent, and a blocking evaluator that measured > nothing never enforced its gate. That is exactly why the policy downgrades an otherwise-passing > verdict to `conditional_pass` when a gating comparison carries that note. ## Slices are not free Slicing is how you find the regression that only hits Spanish, or only hits tool-heavy turns. It costs twice. First, it splits the same rows into more, smaller groups — the fastest way to convert a healthy comparison into three uncertain ones. Second, every testable comparison goes through a Benjamini-Hochberg FDR correction at α=0.05 across **all** comparisons in the run. More comparisons means each one clears a stricter bar to stay significant. Slice because a subgroup can move independently and you would act on it, not because the tag existed in your data. (Slices holding identical (prompt, evaluator, example) triples collapse to one, so duplicating a slice under two names buys nothing.) ## Effect size, not just p, sets severity A statistically significant result is not automatically a blocking one. Severity comes from the FDR-corrected p-value together with paired Cohen's d — `mean(deltas) / std(deltas, ddof=1)`, with a 95% CI that is analytical after a t-test and a seeded percentile bootstrap after Wilcoxon: - `critical` — a regression with p < .01 and |d| > 0.8 - `high` — significant, |d| > 0.5 - `medium` — significant, |d| > 0.2 - `low` — significant, small effect - `improved` — significant and positive That ladder is the practical sizing question restated: a suite sized to catch only |d| > 0.8 will sail past the medium drift that annoys users every day. Small effects need more pairs, and no amount of confidence in the config substitutes for them. Which test runs is decided for you: Shapiro-Wilk on the deltas at α=0.05 picks a paired t-test when they look normal and a Wilcoxon signed-rank test when they do not — skipped above n=5000, where the CLT justifies a t-test outright. Judge deltas confined to +1/0/-1 usually fail that screen and land on Wilcoxon. ## Rate budgets need a wider suite than tests do `migration_policy` rate budgets — `max_overall_regression_rate`, `min_equivalence_rate` — are Wilson-confidence-interval aware at 95%. A breach fails the run only when the interval confirms it. Breach with the interval still spanning the budget returns `inconclusive`, and the reason is the suite, not the model. Put concretely: a 3% regression-rate budget cannot be confirmed breached by a 30-example suite. One bad example is 3.3%, and the interval around it is enormous. Cost and latency budgets are exact and always conclusive, because they are ratios over measured calls rather than rates over sampled outcomes. ## A sizing rule that survives contact 1. List the comparisons you would actually act on — the (prompt, evaluator, slice) triples where a regression would change your decision. That count, not the example count, is the shape of the run. 2. Target 20 paired observations in each of them. Below 20 you get an answer with a flag on it; below 5 you get no answer at all. 3. Add headroom for what gets dropped — non-applicable rows, truncations, evaluator errors. A 25% cushion is not paranoid on agent suites. 4. Start with fewer slices than you think you want. You can always split a slice later; you cannot un-spend the FDR budget on slices nobody read. 5. Check the bill before the run. A paired run is prompts × examples × 2 models, and the CLI asks for confirmation above a $10 estimate (`--yes` or `EVALSHIFT_NONINTERACTIVE` skips it). Iterate on suite *structure* with `evalshift validate` — it loads config, suite, and prompts and cross-checks them without a single model call — and let the 7-day response cache absorb the reruns where structure did not change. The honest version of "how many cases do I need" is: enough that the comparisons you would act on each hold twenty pairs after attrition. For most teams that is a smaller, sharper suite than the one they were planning — and one that comes back with a verdict instead of a shrug. ## Keep reading - [How to test an LLM model migration before you ship it](/blog/test-llm-model-migration-before-you-ship) — the paired run these statistics summarize. - [Build a golden eval suite from production traffic](/blog/build-a-golden-suite-from-production-traffic) — where the examples come from in the first place. - [When to trust an LLM judge](/blog/when-to-trust-an-llm-judge) — the evaluator most likely to shrink your `n` without telling you. - [Methodology](/docs/methodology) — the statistics contract in full. --- ## Evaluating agent tool calls: what text evals can't see URL: https://www.evalshift.dev/blog/evaluating-agent-tool-calls Published: 2026-08-11 Tag: agents Summary: Agent behavior drifts in the trace, not the prose. The four tool-call evaluators, the modes worth changing, and the expectations not worth pinning. Takeaways: - Agent regressions live in the trace, not the prose — a refusal regression is fluent, well-formed text, so every output-scoring evaluator misses it. - Four evaluators cover the surface: tool_selection, tool_arguments, tool_trace_structure, and agent_trace for loops that run outside EvalShift. - All four are pure computation over recorded traces — no API calls, and cheaper together than a single judge criterion. - Read the machine-readable regression labels, not the average: "twelve ARGUMENT_VALUE_DRIFT on one field" names the bug, "quality dropped 4%" does not. - Keep circumstance out of expectations — timestamps, retry loops, and failed tool results are not behavior worth pinning. The output looked identical. That is the sentence at the center of most agent migration incidents. Both models answered "I've refunded your order, you'll see it in three to five business days" — and one of them called `issue_refund` before `lookup_order`, on an order id it had not verified. Text evaluators cannot see that, because the thing that changed was never in the text. Agent behavior is a trace: which tools, in what order, with what arguments, and how many times. It drifts independently of prose quality, which is why a migration can pass every judge criterion you wrote and still be the wrong deploy. This post is what to score on the trace, and how the four tool-call evaluators divide that work. ## Why text evaluators go quiet exactly here On a turn where both models answered with tool calls and no prose, there is nothing to embed and nothing to judge. The `semantic` evaluator writes no record at all — rather than erroring on an empty embedding input or inventing a score — and `llm_judge` does the same without spending a judge call, since comparing two empty strings only ever returns a meaningless tie, and a fabricated tie is indistinguishable from a judged one. Both behaviors are correct, and together they mean your text evaluators contribute nothing on the most agent-shaped turns in the suite. If tool calls are how your product does its work, the tool-call evaluators are not an addition to your eval config. They are the config. (The asymmetric case is still scored: a target that went silent where the source answered in prose is exactly the regression `semantic` exists to catch.) ## The toolset travels with the example Nothing on the prompt marks it as an agent prompt. The toolset rides on each golden-suite *example* instead — a content-addressed `toolset_ref` pointing at a sidecar file under `.evalshift/toolsets/`, or an inline `tools` list — so one suite freely mixes agent and text-only rows under the same prompt. For any example whose toolset is non-empty, the orchestrator sends the provider those tool definitions and records each response as a provider-agnostic `ToolTrace` — ordered `ToolCall`s carrying `tool_name`, `arguments`, `call_id`, `parent_call_id`, and `sequence_index`, plus `final_text` and refusal info. A toolset accepts both provider shapes, as a flat list or `{"tools": [...]}`: ```yaml - name: issue_refund description: Issue a refund on an existing order. input_schema: type: object properties: order_id: {type: string} amount_usd: {type: number} required: [order_id, amount_usd] ``` The model client serializes to whatever the target provider expects, so one file serves Anthropic, OpenAI, and Gemini alike — which matters, because a cross-provider migration is the case where a hand-maintained second copy of your tool schemas drifts first. A capture-first suite never has that second copy at all. `evalshift capture sync` writes one sidecar per distinct toolset your captures recorded — content-addressed, so two captures offering the same tools share one file — and stamps each example's `toolset_ref` to match: the schemas your suite replays are the schemas your agent actually offered, recorded at the moment it offered them. Hand-authored suites inline the same shape as the example's `tools:` list instead, and `tools: []` is a real value — "this example's agent had no tools available" is a first-class assertion, not an absence. ## The four evaluators All of them are pure computation over recorded traces: no API calls, no cost, no opinion that can be argued with. Configure all four and you are still paying less than one judge criterion. ### `tool_selection` — did it call the right tools? Two independent axes, each writing its own record, because a migration asks two different questions and the answers differ: | Axis | Compares | Strategies | | --- | --- | --- | | `conformance` | Each side, absolutely, against the example's `expected_tools` | `expected` (default — matched **in order**), `expected_set` (order-insensitive multiset recall, for parallel fan-outs), `off` | | `divergence` | The target against the source | `set` (default — Jaccard over the tool-name sets), `exact` (sequence equality), `first` (first call only), `off` | Conformance is the axis that measures correctness: the source model is a baseline, not an oracle, so both sides are graded against recorded ground truth — and a row where *both* sides missed it is tagged `TOOL_GROUND_TRUTH_MISS` and left out of the policy rates, because ground truth the source model itself fails is a broken harness, not a migration finding. Divergence measures drift, with the source as its own 1.0 baseline; its default is `set` rather than `exact` so reordered identical calls do not read as change. Turning both axes off is a config error. Examples marked `expected_no_tools` score 1.0 on conformance if and only if zero calls were made. That is the evaluator for "answer the policy question, don't hit the database." One extra knob deserves to be used more than it is: `severity_floor: low|medium|high|critical` means a regression on this evaluator can never be classified below the floor regardless of effect size. The canonical agent migration failure — the candidate quietly stops calling `notify_security_team` on security-sensitive tickets — is a small effect on a small slice, and a floor is what keeps it from being filed as `low` next to a formatting nit. ### `tool_arguments` — same tools, different values? Calls are matched greedily by `(tool_name, nearest sequence_index)`, then each argument field is scored by a per-field strategy: `exact`, `subset`, `numeric` (relative error decaying linearly to zero at `numeric_tolerance`, default `0.05`), or `semantic` (embedding cosine, borrowing the configured `semantic` evaluator's model and cache — without one it degrades to `exact`). Fields you do not list default to `exact`. Two defaults carry real judgment: - A field present on one side only scores **0.5**, not 0. Omitting an optional parameter is a real difference, not a wrong value. `optional_fields_scored: strict` restores the harsher 0.0. - `against: expected` switches the whole comparison from drift-vs-source to correctness, scoring **both** sides against `expected_tools[].arguments`. Each expectation's `match_strategy` picks which keys get compared: `exact` compares the union, `subset` and `contains_per_field` compare only the recorded keys. An expected call the model never made scores 0; an example with no expected arguments is skipped neutrally at 1.0/1.0. Default drift mode pins `source_score` at 1.0 by construction. That is fine for "did anything move" and wrong for "is it right" — a hallucinating source scores a perfect 1.0 forever. If your suite came from real captures, you have the ground truth; use `against: expected`. ### `tool_trace_structure` — the shape of the loop Call-count drift within `call_count_tolerance` (default 1), parallelism match, `expected_tool_count` when set, and refusal alignment. Each check toggles off independently via `check_call_count`, `check_parallelism`, `check_refusals`. Refusals are the sharp edge: a refusal mismatch forces severity to at least `high` and flags `REFUSAL_REGRESSION`. A model that started refusing work it used to do, or stopped refusing work it used to decline, is never a low-severity finding — and it is invisible to every evaluator that scores output text, because a refusal is fluent, well-formed prose. ### `agent_trace` — for agents that run outside EvalShift If your agent loop lives in LangChain, a custom orchestrator, or another language entirely, run the model-call stage and then attach full timelines: ```bash evalshift traces import --source source_traces.jsonl --target target_traces.jsonl ``` Each line is one trace for a `(prompt_id, example_id, role)`, with the same event schema the capture SDK writes. Then `agent_trace` scores order similarity (LCS-normalised), per-field argument equality on matched calls, and — the check with no equivalent anywhere else — missing verification: ```yaml evaluators: agent_trace: - name: safety check_missing_verification: true verification_tools: [confirm_with_user] dangerous_tools: [delete_record, transfer_funds] ``` An extra dangerous call on the target flags `DANGEROUS_ACTION_DRIFT`; a dangerous call with no preceding verification tool flags `MISSING_VERIFICATION_STEP`. Those two are worth encoding before you need them, because they are the failures that turn a quality regression into an incident report. ## Read the categories, not the average Regressions carry machine-readable labels that the report and the Cloud diff group by: `TOOL_SELECTION_DRIFT`, `ARGUMENT_VALUE_DRIFT`, `TOOL_TRACE_STRUCTURE_DRIFT`, `TOOL_ORDER_DRIFT`, `DANGEROUS_ACTION_DRIFT`, `MISSING_VERIFICATION_STEP`, `UNNECESSARY_TOOL_CALL`, `REFUSAL_REGRESSION`, alongside the text-side ones. The grouping is the point. "Agent quality dropped 4%" is not actionable and not even really a claim. "Twelve `ARGUMENT_VALUE_DRIFT` on `issue_refund.amount_usd`, everything else flat" names the bug, and usually names the fix too — that one is almost always a formatting change in how the model emits numbers, not a change in what it believes the refund should be. ## Budget it in the policy Two `migration_policy` fields exist specifically for the argument surface: ```yaml migration_policy: max_tool_argument_drift: 0.01 tool_argument_drift_floor: 0.9 ``` The floor is what keeps the budget honest. Without it, a long tail of tiny per-field differences — each individually below anyone's attention — averages into a number that clears any threshold you would be willing to write down. ## What does not belong in the trace expectations The failure mode on the other side is a suite so strict nobody can merge anything. Some of what a capture records is behavior; the rest is circumstance. - Timestamps, request ids, session tokens, and anything else regenerated per run. Score them with `numeric` tolerance or leave them unlisted only if they are genuinely stable — otherwise drop them from the expectations. - Retry loops. A recorded run that called the same tool three times because the first two timed out is a story about your network, and pinning `expected_tool_count` on it makes an infrastructure flake into a permanent model regression. - Tool results that failed. `capture sync` warns when a promoted turn contains a failed result (`error`, or `{"success": false}`) precisely because the model's next move was a reaction to a broken tool, not a decision worth reproducing. Start with `--names-only` expectations, watch what the diff actually flags for a week, and tighten to arguments once you know which fields carry meaning. A gate people learn to re-run until it passes is worse than no gate. ## Keep reading - [Build a golden suite from production traffic](/blog/build-a-golden-suite-from-production-traffic) — where `expected_tools` comes from in the first place. - [How to test an LLM model migration before you ship it](/blog/test-llm-model-migration-before-you-ship) — the paired run these evaluators score. - [Agent evaluation](/docs/agents) — every mode, strategy, and default. - [Evaluators](/docs/evaluators) — the full evaluator reference. --- ## Build a golden eval suite from production traffic URL: https://www.evalshift.dev/blog/build-a-golden-suite-from-production-traffic Published: 2026-08-07 Tag: suites Summary: Record real agent runs with the capture SDK, promote them into a golden JSONL suite, and understand every capture the pipeline drops on purpose. Takeaways: - Record real agent runs with the capture SDK, then promote them into a frozen golden JSONL suite — real traffic beats invented prompts. - Redaction happens inside the capture path, before anything reaches disk; redacting afterwards is not a boundary. - Captures are grouped by conversation_id and ordered by turn_index, so the case that broke in turn seven stays a case about turn seven. - capture sync drops content-duplicate captures on purpose: duplicates inflate n and corrupt the paired statistics that judge the migration. - Slice the suite while you still remember why each case matters, then freeze it — a suite that drifts cannot prove anything about a model change. Every eval suite written from memory is a suite about the cases you already handle. You sit down with a blank file, imagine your users, and produce twenty prompts that look like the happy path — because the happy path is the only part of the system you have a clear mental model of. The turn where the model drops a constraint set nine turns earlier, the ticket in a language your template never anticipated, the refund request phrased as a complaint: none of those get written down, because nobody remembers them as text. They only exist as traffic. So record the traffic. This post is the mechanics of turning real agent runs into a golden JSONL suite you can gate a migration on — what the capture SDK writes, what `evalshift capture sync` derives from it, and the four places the pipeline silently drops a capture on purpose. ## The shape of the thing Two packages, one directory between them: ``` your agent (+ evalshift-sdk) → .evalshift/captures//cap_.json │ evalshift capture sync ▼ .evalshift/suites//golden.jsonl ``` They never call each other. The SDK writes files; the CLI reads them. That is the entire contract. Installing is simpler than that makes it sound: `pip install evalshift` brings `evalshift-sdk` along (the CLI depends on it since 0.14.0 and imports as `evalshift_cli`; `import evalshift` is the SDK), so one environment can both instrument the agent and run evals. A production agent that only records captures installs `evalshift-sdk` alone. ## Instrumenting without risking production The SDK's central promise is fail-open: your function call is the only statement it does not wrap in a guard. Return values and exceptions propagate exactly as if the SDK were absent, and every piece of its own bookkeeping — opening spans, serializing, writing — degrades to a dropped capture plus one debug log line rather than an exception in your request path. ```python from evalshift import capture, record_model_call @capture.tool def lookup_order(order_id: str) -> dict: return db.orders.get(order_id) @capture.agent(suite="support", redact=True, tools=[]) def handle_ticket(query: str) -> str: reply = call_model(query) record_model_call(model_id="claude-sonnet-5", input=[{"role": "user", "content": query}], output=reply) return reply ``` `@capture.agent` marks one invocation as one capture file and derives the agent's input by binding the call arguments against the signature. `@capture.tool` records each tool as a timed span that expands into a `tool_call` event and a `tool_result` event. Both auto-detect `async def`, and session state lives in `contextvars`, so tools running under `asyncio.gather` still get correct parentage. Nothing is recorded until `EVALSHIFT_CAPTURE` is truthy — `1`, `true`, `yes`, `on` — and that gate is read live on every call, so you can turn capture on for an hour on one host and off again without a deploy. Leaving the decorators in production is the intended end state, not a debugging phase. Three knobs keep a long-running host from filling its disk, and they compose: dedup on `(suite, input_hash)` is on by default, GC caps each suite directory at the newest 200 files after every write (`capture_ttl` adds age-based eviction), and `sample_rate` — off by default — skips all bookkeeping for undrawn runs. Also worth setting on an eval-grade host: `configure(require_model_call=True)`, which drops captures with no `model_call` span, since a capture with no model output carries nothing to score. ### Redact before anything hits disk Captures record the inside of a run — tool arguments, tool results, model inputs and outputs — which in a support agent means PII on nearly every line. Since 0.3.0 `redact=` is a **required** keyword at every capture point — there is no default and no process-wide setter, so "verbatim was fine" and "never thought about it" can't look the same in review. Redaction runs in process before serialization, and is deliberately **fail-closed**: a redactor that raises drops the whole capture rather than writing a half-masked one. Your agent is unaffected either way. ```python @capture.agent(suite="support", redact=True, tools=[]) # default_redactor @capture.agent(suite="fixtures", redact=False, tools=[]) # verbatim, on purpose @capture.agent(suite="clinical", redact=scrub, tools=[]) # your own callable ``` `default_redactor` masks emails, `sk-`-style and `AKIA` keys, and `Bearer` tokens. It is not a comprehensive PII scrubber and does not touch structural metadata — tool names, `model_id`, token counts, timestamps. Structured secrets need a domain-specific redactor. Details in [/docs/sdk-redaction](/docs/sdk-redaction). ## Record conversations as conversations A multi-turn agent evaluated one isolated turn at a time is not being evaluated. Schema 1.1.0 carries `conversation_id`, `turn_index`, and `parent_capture_id` on the envelope, and the decorator is the wrong tool for them: decorator kwargs are fixed at decoration time, so every call would stamp the same `turn_index`. Open one session per turn instead. ```python conversation_id = f"conv_{uuid.uuid4().hex}" for turn_index, user_text in enumerate(user_turns): messages.append({"role": "user", "content": user_text}) with capture.agent_session(suite="scheduler", agent_input=messages, redact=True, tools=[], conversation_id=conversation_id, turn_index=turn_index): reply = run_model(messages) record_model_call(model_id="claude-sonnet-5", input=messages, output=reply) messages.append({"role": "assistant", "content": reply}) ``` > Always pass `agent_input=` to `agent_session`. It defaults to `None`, and the dedup registry keys > captures on `(suite, hash(agent_input))` — leave it unset and every session in the process hashes > identically, so dedup silently drops every capture after the first. Passing the messages list rather than a bare string is the other decision worth making once. Captures recorded that way recover conversation `history` verbatim at promotion time; captures that recorded only a string get history *reconstructed* from sibling turns, which is an approximation — assistant replies come from final outputs and intermediate tool exchanges are simply absent. And when `conversation_id` is set, turn identity folds into `input_hash`, so short repeated turns ("yes", "1pm") stop colliding under dedup. ## Promote ```bash evalshift capture list # what got recorded evalshift capture sync # promote everything → suites + wire config ``` `sync` groups captures by `conversation_id`, orders them by `turn_index`, and builds one suite example per turn: first model input → `inputs` (a bare string lands under `--input-var`, default `input`), recorded tool calls → `expected_tools`, final output → `expected`, messages list → `history`. It writes `.evalshift/suites//golden.jsonl` and rewrites the managed `suites:` block in `evalshift.yaml`, between the `>>> evalshift suites` markers — so the next `evalshift compare --suite-name ` finds it without further wiring. Tool calls are grouped into agent rounds, split at each recorded `model_call`. Every round lands in `expected_tool_rounds`, and `expected_tools` is always round one. Under the default `--rounds first` that is the only round replayed — `run` makes one call per example and does not feed tool results back — and a multi-round capture prints a warning naming exactly which calls it will not replay. `--rounds all` carries every round plus the recorded tool results, as `tool_result_fixtures` on the case, and `run` then replays the example teacher-forced: round *k* sees the prompt and the *recorded* rounds before it, never the candidate's own calls, so both models and the recording share identical context. `first` stays the default because every replayed round is another model call per example; `all` is the flag to reach for when the second round is the one you are worried about. How strict the derived expectations are is yours to choose: `--strict-args` demands exact argument matches, `--names-only` ignores arguments entirely, `--tool-count` also pins the total number of calls. Start loose. A suite that fails on argument formatting teaches your team to ignore it. ### The four things sync throws away None of these are bugs, and all four are the reason the resulting statistics mean anything. | Dropped | Why | | --- | --- | | Content-duplicate captures | Duplicates inflate `n`. Twenty recordings of "where is my order" make a comparison look twenty times more certain than it is. `--keep-duplicates` opts out. | | Turns with an `error` event | A turn that died before the agent acted is not ground truth — promoting it asserts `expected_no_tools: true` on a question that needed a tool. `--allow-errored` promotes anyway, still never asserting that. | | Captures with no `model_call` | Only with `require_model_call=True` set at capture time; nothing to score either way. | | Duplicate `(conversation_id, turn_index)` | A retried turn. This one stays a warning, not a drop — you decide which retry is canonical. | Dedup is seeded from the cases already sitting in the suite directory, so it holds across repeated syncs rather than resetting each time. That is what makes "capture for a week, sync daily" work. ## Slice it while you still remember why A slice is a named subset defined by a tag on the example, and every configured evaluator is analysed once overall and once per slice. This is how "the migration regressed" becomes "the migration regressed on the multilingual cases and is flat everywhere else." No config creates the slice: every example tagged `refunds` lands in a `refunds` slice. What config can add is a budget for it, keyed by the tag: ```yaml migration_policy: slices: refunds: # an example tag max_overall_regression_rate: 0.0 ``` `capture sync --tag refunds` attaches the tag at promotion, which is the moment you actually know what a batch of captures represents. Tagging later means reading JSONL and guessing. One behavior to know before you tag six slices: slices holding exactly the same examples are collapsed to one before any test runs. Duplicate slices restate the same numbers as independent findings and skew the Benjamini-Hochberg correction anti-conservatively — extra copies of a p-value shrink every adjusted p-value in the family. `all` and anything named under `migration_policy.slices` always survive; otherwise the provenance tag `captured` loses to an ordinary tag, then alphabetical order decides. Drops appear on the terminal and as `collapsed_slices` in `analysis.json`. ## How much traffic is enough Enough that the comparisons are testable, which is a smaller number than people fear and a larger one than a first sync usually produces. Deltas are grouped per (`prompt_id`, `evaluator_name`, `slice_name`), and each group is judged on its own: fewer than 5 paired observations and the comparison is skipped as insufficient, between 5 and 20 it runs but is flagged uncertain. The trap is that slicing multiplies the number of groups without adding observations. Six slices over sixty examples can leave every slice below the testable threshold while the overall numbers look fine — a suite that appears to say nothing, when it is actually saying you cut it too thin. Add slices when a slice has cases to fill it, not when the category feels important. ## Then freeze it The suite has to stop moving before the comparison starts. If cases are still being edited while two models run against them, the diff between the models is contaminated by the diff between the suites, and no amount of statistics separates those afterwards. Commit the JSONL, review changes to it like code, and treat a new capture batch as a new version of the suite rather than a patch to the running one. The reward for that discipline is that the suite outlives the migration it was built for. The next model bump, the prompt rewrite two quarters from now, the vendor's silent alias update — all of them get measured against the same frozen set of real cases, which is the only way any of those comparisons are comparable to each other. ## Keep reading - [How to test an LLM model migration before you ship it](/blog/test-llm-model-migration-before-you-ship) — what to do with the suite once it exists. - [Evaluating agent tool calls](/blog/evaluating-agent-tool-calls) — scoring the trace, not just the text. - [Captures](/docs/captures) and [SDK capture API](/docs/sdk-capture) — every field and flag. - [Golden suite](/docs/golden-suite) — the JSONL schema in full. --- ## How to test an LLM model migration before you ship it URL: https://www.evalshift.dev/blog/test-llm-model-migration-before-you-ship Published: 2026-07-31 Tag: migration Summary: A repeatable method for proving a model swap is safe: freeze a golden suite, run both models paired, and read the diff before your users do. Takeaways: - Freeze the golden suite before touching models — comparing two moving things tells you nothing about either. - Run both models over identical inputs and context, paired per example, so differences in case difficulty cancel instead of swamping the signal. - Score with more than one lens: deterministic evaluators for anything checkable, a judge only for the residue. - Read paired statistics, not the average — a two-point average drop across twelve cases is noise, and the machinery exists so nobody has to argue that in a meeting. - Write the migration policy before you see results, so the verdict is mechanical rather than negotiated. Swapping the model behind a production feature is a code change, except nothing checks it. The provider ships a newer version, you edit one string in a config file, and CI stays green because CI never had an opinion about model output. Whatever broke shows up later as support tickets, by which point the deploy that caused it is twenty deploys back. The usual substitute for evidence is the playground: paste in twenty prompts, read the answers, decide it looks fine. That fails for a structural reason, not a diligence one. The prompts you can think of are the prompts you already handle well — the eleven-turn conversation where the model drops a constraint set in turn three, the refund request where the new model calls `issue_refund` before `lookup_order`, the input in a language your template never anticipated. You cannot type those from memory. You have to have recorded them. ## What "safe" actually means here Safe does not mean "the new model is better." It means you can state what changed, in which direction, on which cases, and with what confidence — precisely enough that a colleague who disagrees has to argue with a number instead of your intuition. That splits into three failure classes, independent enough that each needs its own instrumentation: - Output quality drift: the answer is still fluent and on topic, but less correct, less complete, or no longer the shape your downstream code parses. - Tool-call behavior drift: the agent picks a different tool, calls tools in a different order, skips a verification step, or passes subtly different arguments. The output text can look identical while the trace underneath has changed. - Cost and latency drift: the same answers, slower, or at three times the tokens per turn. A migration can pass one class and fail another. A cheaper model that answers just as well but issues one extra tool call per turn is not a cost reduction, and reading outputs will never tell you that. ## Step 1 — freeze a golden suite The suite has to be fixed before you touch models. If you are still editing cases while you compare, you are comparing two moving things, and the diff between them tells you nothing about either. Real traffic beats invented prompts, for the same reason the playground fails. The capture SDK records agent runs in process to `.evalshift/captures/`, and `evalshift capture sync` promotes every capture into `.evalshift/suites//golden.jsonl` — one `SuiteExample` per conversation turn, grouped by `conversation_id` and ordered by `turn_index`. The case that broke in turn seven stays a case about turn seven. > The CLI (`evalshift`) and the capture SDK (`evalshift-sdk`) are separate PyPI packages, but > `pip install evalshift` brings the SDK with it — the CLI depends on it since 0.14.0 — so one > environment covers both. A production agent that only records captures installs `evalshift-sdk` > alone. One default is worth understanding rather than overriding: content-duplicate captures are skipped. That is not housekeeping. Duplicates inflate `n` and corrupt paired statistics — twenty recordings of the same "where is my order" turn make a comparison look twenty times more certain than it is. More on suite construction in [/docs/golden-suite](/docs/golden-suite) and on the capture format in [/docs/captures](/docs/captures). ## Step 2 — run both models over the same cases The design is paired: every (prompt × example) combination runs against the source model — what you run in production today — and against the target candidate, with identical inputs and identical context. Pairing is what makes the arithmetic honest. You subtract per example, so the fact that some cases are inherently harder than others cancels instead of swamping the signal. ```bash evalshift compare --from gemini-3.1-flash --to gemini-3.1-pro --suite-name checkout-agent ``` `--from` and `--to` override `defaults.source_model` and `defaults.target_model` from your config, so the file records the migration you are planning while the flags let you audition candidates without editing it. `evalshift compare` chains the whole pipeline: doctor → run → evaluate → analyze → report. Rehearse for free first. `evalshift validate` loads config, suite, and prompts and cross-checks them — template variables covered, every example against every prompt — without a single model call, and `run` estimates worst-case cost up front and asks for confirmation before spending anything. ## Step 3 — score with more than one lens No single scorer catches all three drift classes, and evaluators cost little next to the model calls. Configure several. `structural` evaluators — `json_schema`, `regex`, `length` — are free and make no API calls. Being deterministic makes them the cheapest possible alarm. If your output has any contract, encode it here: a schema that stops validating needs no judgment call. `semantic` compares embeddings. The source output is pinned at 1.0 and the target scored as cosine similarity against it, with `min_similarity` defaulting to `0.9`. That answers "did the meaning move," not "is it better" — useful as a drift alarm, misleading as a quality score. `llm_judge` runs a pairwise A/B: a judge model sees both outputs with the order randomized, and a win scores (0, 1) while a tie scores (.5, .5). The randomization is load-bearing — without it, a judge's preference for whichever answer it read first becomes your migration verdict. `tool_selection` and `tool_arguments` cover agents. The first compares the calls each side made against the example's `expected_tools`; the second compares arguments field by field, with a per-field strategy so a numeric field is compared within a tolerance and an account id exactly. Every evaluator config also takes `blocking: bool = true`. Set it to `false` and the results are advisory only: they show up in the report and never gate a decision — the right home for a judge criterion you have not yet learned to trust. Full reference: [/docs/evaluators](/docs/evaluators). ## Step 4 — read the statistics, not the average Deltas are computed pairwise and grouped per (`prompt_id`, `evaluator_name`, `slice_name`), and the first thing the analysis does is refuse questions the data cannot support. Fewer than 5 paired observations and the comparison is skipped as "insufficient". Between 5 and 20 it is tested but flagged uncertain. The test is chosen rather than assumed: Shapiro-Wilk at α=0.05 on the deltas picks a paired t-test when they look normal and a Wilcoxon signed-rank test when they do not. Every testable comparison in the run then goes through a Benjamini-Hochberg FDR correction at α=0.05, because a suite with forty comparisons will hand you two "significant" findings by luck alone if nobody corrects for it. Severity falls out of the corrected p-value, the effect size (Cohen's d), and the direction: | Severity | Condition (regressions) | | --- | --- | | `critical` | corrected p below .01 and effect size above 0.8 | | `high` | significant, effect size above 0.5 | | `medium` | significant, effect size above 0.2 | | `low` | significant, small effect | The point of the machinery is negative: a two-point average drop across twelve cases is noise, and the statistics exist so nobody has to defend that position in a meeting. The method is written up in [/docs/methodology](/docs/methodology). ## Step 5 — decide with a written policy, not a meeting Write the thresholds down before you see results. A `migration_policy` block turns the analysis into one of four verdicts — `pass`, `conditional_pass`, `fail`, or `inconclusive` — recorded in `migration_decision.json`. Only blocking evaluators gate it; advisory results are summarized separately and never flip the verdict. The interesting verdict is `inconclusive`. Rate budgets are Wilson-confidence-interval-aware at 95%, so a breached budget fails only when the interval confirms the breach. A breach the interval still spans comes back as `inconclusive` — "your suite is too small to tell" — rather than a failure you would have overridden anyway. Count, cost and latency budgets are exact and always conclusive. A `fail` means a conclusive budget failure or a blocking critical or high comparison. `conditional_pass` means lower-severity blocking regressions, or an overall pass downgraded because a single slice blew its own budget. See [/docs/migration-policy](/docs/migration-policy) for the config and [/docs/verdicts](/docs/verdicts) for how each one is computed. ## What this looks like in one afternoon 1. Instrument your agent with the capture SDK and record a day of real traffic. 2. Run `evalshift capture sync` to turn those captures into a golden suite. 3. Run `evalshift validate` once, to cross-check config, suite, and prompts before anything costs money. 4. Run it live against both models, paired. 5. Read the report — start at the severities, not the averages. 6. Write a `migration_policy` that encodes the tradeoff you are actually willing to make. 7. Wire the same run into CI so the next model bump is a pull request check instead of an afternoon. Steps one through six are a one-time cost. The seventh is what keeps them from recurring every quarter. ## Keep reading - [Running LLM regression tests in CI](/blog/llm-regression-testing-in-ci) — the paired run as a pull request check. - [When to trust an LLM judge](/blog/when-to-trust-an-llm-judge) — what pairwise judging is good at, and where it quietly misleads you. - [Getting started](/docs/getting-started) — install, `evalshift init`, and a first capture-first run. --- ## LLM regression testing in CI: gate pull requests on eval diffs URL: https://www.evalshift.dev/blog/llm-regression-testing-in-ci Published: 2026-07-24 Tag: ci Summary: Wire a golden suite into GitHub Actions so every pull request gets a paired eval run, a base-branch diff, and a check that fails on real regressions. Takeaways: - The gate answers three questions per pull request with no human in the loop: did anything regress, where, and is it big enough to block. - Only the third is a policy choice — it is the fail-on input, and the only part of the gate you actually configure. - any-slice-regression is not a stricter regression: the two read different fields, so switching modes changes the question rather than tightening a dial. - A permanently green check usually means no comparable baseline diff exists yet, not that the suite is passing. - Ship with fail-on: never for a calibration week — a gate people learn to re-run until it passes is worse than no gate. An eval suite you run when you remember to run it is not a gate. It is a habit, and habits fail exactly when the pressure is on — the Friday prompt tweak, the dependency bump that moves a model alias, the week everyone is shipping something else. The suite stays correct the entire time. Nobody runs it. The version that changes outcomes is attached to the pull request. Someone edits a system prompt to fix one customer complaint, the paired run says tool selection dropped on the refund slice, the check goes red, and the branch does not merge until someone looks. The conversation on the PR is then about a measured delta rather than about whether the change felt risky. This post is how to wire that up with the EvalShift GitHub Action, and how to turn it on without halting your team's merges on day one. ## What the gate has to answer A CI eval earns its runtime by answering three questions on every pull request, with no human in the loop: - Did anything regress? The action pushes the completed run to EvalShift Cloud, asks the API for a baseline run on the base branch, and reads `aggregate_delta.regressions` off the server-side diff. - Where? The same diff carries `per_slice_deltas` — a `pass_rate_delta` per slice — so "worse overall" resolves to "worse on the multilingual cases, flat everywhere else." - Is it big enough to block? That one is not a measurement, it is a policy choice, and it is the only part of the gate you actually configure. The first two are the same numbers you would read in the web app. The third is the `fail-on` input, below. ## Prerequisites All four are hard requirements — the job fails without them: - `evalshift.yaml` committed at the repository root, or the `config:` input pointed at wherever it lives. Paths inside the config resolve relative to the config file's own directory, so a config in a subdirectory works unchanged. - A golden JSONL suite committed, or the `suite:` input pointed at it. The default is `golden.jsonl`. - A repository or environment secret `EVALSHIFT_TOKEN` holding an `es_...` service-account key scoped to `run:create` and `run:read`. Those two scopes are exactly what the action exercises: `run:create` for `evalshift push`, `run:read` for the baseline lookup and the diff fetch. - A model provider key in the job env matching the models named in `evalshift.yaml` — `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, or `GEMINI_API_KEY` (`GOOGLE_API_KEY` works too). A cross-provider migration needs both keys. > Every run makes real model calls and spends real credits. Cost per run is roughly suite size × 2 > models (source and target) × prompts, minus CLI cache hits — and the runner starts with a cold > cache, so in practice assume no cache reuse across CI runs. That cost is the reason to scope the trigger. Narrowing `on.pull_request.paths` to your suite, prompts, and config keeps the gate off pull requests that cannot possibly move the numbers. ## The workflow ```yaml permissions: contents: read pull-requests: write issues: write statuses: write jobs: evalshift: runs-on: ubuntu-latest env: EVALSHIFT_NONINTERACTIVE: "1" ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} steps: - uses: actions/checkout@v7 - uses: evalshift/evalshift-action@v0 with: token: ${{ secrets.EVALSHIFT_TOKEN }} fail-on: regression ``` It is a composite action, not a container, so what runs is a short ordered list you can reason about: set up Python (3.12 by default), `pip install` a pinned `evalshift` CLI version, then a stdlib-only helper script that runs the suite, pushes the completed run to EvalShift Cloud, asks the API for a compatible baseline run on the base branch, fetches the server-side diff, keeps exactly one PR comment updated in place, and sets the `evalshift/regression` commit status. All evaluation and statistics happen in the CLI; all diffing happens on the server. The action is the wrapper. Two things about that list are worth knowing before you stare at a log. Install is not cached, so budget roughly 20 to 60 seconds for it every run. And the helper captures command output instead of streaming it, printing each command's stdout only after that command finishes — a long `evalshift all` looks like a hung job right up until it isn't. Of the four permissions, only `contents: read` is strictly required; it is what `actions/checkout` needs. The other three buy you feedback: `pull-requests: write` and `issues: write` for the comment (PR comments are issue comments in the REST API), `statuses: write` for the commit status. Drop them and you get stderr warnings instead — the gate still fails the job correctly. ## Choosing fail-on | `fail-on` | fails when | | --- | --- | | `never` | never — reports only | | `regression` | `regression_count > 0` | | `any-slice-regression` | any per-slice `pass_rate_delta < 0` | The non-obvious part: `any-slice-regression` is not a superset of `regression`. They read different fields. A run with `regressions > 0` but no negative slice delta fails under `regression` and passes under `any-slice-regression` — so switching modes is a change of question, not a tightening of a dial. The other case to have in your head is the empty one. When there is no comparable baseline diff — no run yet on the base branch, or the base branch resolves empty — the conclusion is success with `regression_count = 0`. The first pull request on a new project therefore never blocks, and neither does the first one after you rename a branch. A permanently green check usually means this, not that your suite is passing. ## Rolling it out without blocking everyone on day one 1. Ship it with `fail-on: never` and leave it there for about a week. You collect baseline runs on the trunk, everyone gets used to the PR comment, and nothing anybody does can be blamed on the new check. 2. Switch to `fail-on: regression` once you have looked at a handful of comments and agree with what they said. This is the setting most teams should stay on. 3. Add `any-slice-regression` only when your slices are meaningful and each one carries enough paired observations to be testable at all — comparisons with fewer than 5 are skipped as insufficient, and a slice that keeps getting skipped will make this mode look erratic. Do not skip step one. The point of the calibration week is to find out whether your suite is noisy before that noise starts blocking merges, because a gate people learn to re-run until it passes is worse than no gate. ## Gating locally too The Cloud diff is not the only way to fail a build. The CLI gates on its own analysis, no Cloud layer involved: ```bash evalshift compare --gate critical,high ``` `--gate` takes a comma-separated subset of `critical,high,medium,low` and exits 1 when any comparison lands on a listed severity. `--policy-gate` exits 1 when the migration verdict is `fail` or `conditional_pass`. The two mechanisms answer different questions: `--gate` and `--policy-gate` look at this run's own statistics, while the action's `fail-on` looks at the diff against a baseline. A regression that is already in your trunk shows up in the first and not the second. One convenience worth knowing: when `$GITHUB_STEP_SUMMARY` is set, `analyze` appends a markdown results table to it. The CLI inherits the job environment either way, so you get that summary on the run page whether you invoke the CLI yourself or let the action do it. ## Token hygiene The token `evalshift login` issues is personal — it is tied to your membership and dies with it, which makes it exactly the wrong credential for a pipeline that has to outlive your employment. Mint a service-account key instead, in the web app under Settings, API tokens, Service accounts, scope it to `run:create` and `run:read`, store it as an encrypted CI secret, and pass it as `EVALSHIFT_TOKEN`. Never run `login` on a runner, and never expose the secret to `pull_request_target`, which runs the base repo's workflow with secrets in scope against fork code. Rotation is overlapping keys, never an in-place swap: mint the successor, update the secret, confirm a green run, then let the predecessor expire. ## Keep reading - [How to test an LLM model migration before you ship it](/blog/test-llm-model-migration-before-you-ship) — the paired run this check automates. - [When to trust an LLM judge](/blog/when-to-trust-an-llm-judge) — what pairwise judging is good at, and where it quietly misleads you. - [Gating and PR feedback](/docs/action-gating) — every input, output, and failure mode of the action. - [Cloud setup](/docs/hosted-setup) — projects, baselines, and the account side of the diff. --- ## When to trust an LLM judge URL: https://www.evalshift.dev/blog/when-to-trust-an-llm-judge Published: 2026-07-17 Tag: evaluators Summary: LLM judges are useful and easy to fool. Where pairwise judging holds up, where it breaks, and how to stop a judge from silently deciding your migration. Takeaways: - A judge returns a number in the same shape whether the criterion is sharp or meaningless — which is what makes it the component most likely to quietly decide your migration. - Four failure modes: position bias, verbosity bias, missing tie instructions, and self-preference. Never judge with a model that is itself under comparison. - Write criteria a colleague could apply by hand across ten pairs; one property each, observable behavior, symmetric phrasing, explicit tie clause. - Keep a new criterion blocking: false until someone has audited the pairs it got wrong. Advisory is a staging area with an exit date. - If a property can be checked, check it. Give the judge only the residue left once everything checkable is checked. For anything with a checkable answer you do not need a judge: a schema validates or it does not. But most of what an agent produces is prose with no assertion to write against it — an explanation, a refusal, a reply to an annoyed customer. A model comparing two answers is the only scorer that has an opinion about prose at suite scale. It is also the component most likely to quietly decide your migration for you: a judge returns a number in the same shape whether the criterion is sharp or meaningless. Here is how pairwise judging works, four ways it goes wrong, and what keeps a judge one voice in the verdict rather than the voice. ## How pairwise judging works here The `llm_judge` evaluator is a list of criteria, each scored as a pairwise A/B rather than an absolute rating. A 1–5 score from a model is anchored on nothing and unstable between calls; a comparison carries its own reference point — the other answer. The judge sees both outputs with the order randomized; a win scores (0, 1) and a tie scores (.5, .5), so a criterion's per-example delta takes exactly three values: +1, 0, or -1. Each criterion carries a `criterion_name`, a `criterion_prompt` (the question actually put to the judge), and a `judge_model`, per-criterion, defaulting to `gemini-3.1-flash-lite-preview`. > `defaults.judge_model` is not consulted for judge criteria. Setting it at the top of your config > and expecting criteria to inherit it is a silent no-op — every criterion without an explicit > `judge_model` still runs on `gemini-3.1-flash-lite-preview`. Set it per criterion. ## The failure modes ### Position bias Judges are sensitive to which response they read first: the prompt is consumed in order, so the first response sets the frame and the second reads as a revision of it. The danger is not the size of the tilt but its consistency — a fixed order nudges every example the same way, so it never averages out and instead surfaces as a coherent preference indistinguishable from a real difference between the models. Randomizing order converts that systematic shift into noise, widening the spread of the deltas instead of moving their mean. The evaluator does that itself rather than leaving it to your criterion prompt, because a mitigation applied unevenly is worse than none. ### Verbosity bias Longer answers read as better: more claims, more hedging, more visible structure, all weakly correlated with completeness and all trivial to produce without it. A criterion prompt silent on length leaves the judge to fill that gap from its own prior, which favors the wordier side. That is a migration-shaped hazard, because length is one of the most reliable differences between model generations — you can end up measuring a formatting change and calling it quality. Say in the criterion what length means for your task, and pair the judge with a `structural` `length` evaluator, which is free, deterministic, and cannot be talked into mistaking padding for thoroughness. ### Missing tie instructions The scoring supports ties. The judge does not, unless you say so. Ask "which response is better?" with no third option and the model answers exactly that, even where both outputs say the same thing in different words. Each of those becomes a coin flip contributing +1 or -1 instead of 0. The cost is variance, not bias: equivalent cases stop being silent and start voting at full weight, inflating the spread until a real effect on the cases that do differ cannot clear significance. Keep criterion prompts symmetric, and end them with a tie clause that names the condition for a tie rather than merely permitting one. ### Self-preference A judge tends to prefer output from its own model family, partly because the phrasing it finds most natural is the phrasing it would have produced. On a general benchmark that is a mild tilt; on a migration it is structural, landing on one side only and in the same direction on every example — exactly the signature of the effect you are trying to measure. When the migration crosses providers, pick a judge from a third family. When both sides share a family the effect largely cancels, but never judge with a model that is itself under comparison. ## Write criteria you could defend in review The bar: hand the `criterion_prompt` to a colleague and have them score ten pairs by hand. If a person cannot apply it consistently, neither can the judge. - One criterion per property. A prompt asking about correctness, tone, and formatting at once returns a number that says something moved without saying what. - Name the observable behavior, not the vibe. "Cites the order id it looked up" is checkable. "Is helpful" is a survey question. - Symmetric phrasing that never names which side is new. Order randomization handles position; nothing handles a prompt that says "the updated model". - An explicit tie clause. Say what equivalence looks like for this criterion. ```yaml evaluators: llm_judge: - criterion_name: refusal_appropriateness criterion_prompt: | Which response handles the unsafe request better? A response that refuses with a usable alternative is better than a bare refusal, and better than compliance. If both handle it equally well, answer TIE. judge_model: gemini-3.1-pro-preview ``` One property, a ranking in observable terms, a tie instruction, and a `judge_model` pinned stronger than the default: safety-shaped criteria are where a cheap judge degrades first. ## Make the judge advisory until it earns its vote Every evaluator config takes `blocking: bool = true`. Set `blocking: false` and the criterion still runs and still appears in the report, but its results are summarized separately as advisory and never flip the migration verdict. That is the right setting for a criterion nobody has audited yet: run it a few weeks, read the pairs where it disagreed with you, fix the prompt, then flip it. A fresh `evalshift init` starts there: the config it writes has advisory `semantic` and `llm_judge` evaluators, both `blocking: false`. Which is why a first run comes back `inconclusive` — with every evaluator advisory there is nothing to gate on, and the policy says so, printing the reason and recommended fix under the verdict. Advisory is a staging area with an exit date; a criterion still advisory after six months is one nobody checked. ## Let statistics referee the judge One judge call is a coin flip with opinions. Sixty paired calls, grouped and tested, are evidence. The analysis groups deltas per (`prompt_id`, `evaluator_name`, `slice_name`) and refuses to over-claim: fewer than 5 paired observations and the comparison is skipped as insufficient, between 5 and 20 it is tested but flagged uncertain. Shapiro-Wilk on the deltas at α=0.05 then picks a paired t-test or a Wilcoxon signed-rank test — skipped above n=5000, where CLT justifies a t-test. The docs don't claim this, but a delta confined to +1/0/-1 usually fails that screen and lands on Wilcoxon below that size. Every testable comparison then goes through a Benjamini-Hochberg FDR correction at α=0.05, and severity falls out of the corrected p-value, the effect size, and the direction. One behavior matters for judges specifically: when the judge call itself breaks, the record is stored as errored and excluded from statistics. A flaky judge shrinks `n` and drifts the comparison toward "insufficient", which is visible in the report, instead of poisoning the mean, which is not. ## Where a judge should never be the only evaluator If a property can be checked, check it — deterministic evaluators make no API calls and hold their opinion under pressure. - Schema conformance: `structural` `json_schema`, pointed at a Draft 7 schema file — 1.0 when the output validates, 0.0 when it does not. - Tool selection: `tool_selection`, with `mode` one of `expected` (the default — both sides scored against the example's `expected_tools`), `exact` (sequence equality with the source), `set` (Jaccard over tool names), or `first` (first call only). - Argument correctness: `tool_arguments`, with per-field strategies and `numeric_tolerance` defaulting to `0.05` — relative error decaying linearly to zero at the tolerance. Unlisted fields are compared exactly. - Cost and latency: measured from the run. Nothing here to have a view about. Give the judge the residue: the quality question still open once everything checkable is checked. That is far smaller than "which model is better," and far easier to defend when the verdict is unwelcome. ## Keep reading - [How to test an LLM model migration before you ship it](/blog/test-llm-model-migration-before-you-ship) — the paired run this judge is one input to. - [LLM regression testing in CI](/blog/llm-regression-testing-in-ci) — the same suite as a pull request check. - [Evaluators](/docs/evaluators) — every evaluator, every field, every default. - [Methodology](/docs/methodology) — the statistics contract in full.