<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>EvalShift blog</title>
    <link>https://www.evalshift.dev/blog</link>
    <description>Field notes on migrating LLM models safely: golden suites, paired evaluation, regression gates in CI, and evaluator reliability.</description>
    <language>en</language>
    <atom:link href="https://www.evalshift.dev/rss.xml" rel="self" type="application/rss+xml"/>
    <item>
      <title>This one passed: 71% cheaper, and the agent still does the same thing</title>
      <link>https://www.evalshift.dev/blog/what-a-passing-migration-proves</link>
      <guid isPermaLink="true">https://www.evalshift.dev/blog/what-a-passing-migration-proves</guid>
      <pubDate>Wed, 16 Sep 2026 00:00:00 GMT</pubDate>
      <description>An EvalShift migration report that said PASS: a 71% cheaper model inside every policy budget. Where each threshold came from, and what PASS does not claim.</description>
      <content:encoded><![CDATA[The previous post was about a migration that was cheaper, faster and still failed. This is the
other outcome. It is the more common one once a suite is in decent shape, and I see it written up
far less often, because a pass is boring.

It shouldn't be. A pass is only worth something if the limits were fixed before the run. So this
post is the whole thing: the policy, the reasoning behind every number in it, and the EvalShift
report it produced.

For anyone arriving here cold: EvalShift is a local-first CLI for testing LLM model migrations.
It replays a frozen golden suite against your current model and a candidate, scores every pair of
outputs with tool-call, structural, semantic and LLM-judge evaluators, runs paired statistics over
the deltas, and turns a `migration_policy` block you wrote in `evalshift.yaml` into one of four
verdicts: `pass`, `conditional_pass`, `fail` or `inconclusive`. It writes a single-file
`report.html` on your machine; nothing is uploaded unless you run `evalshift push`. Every figure in
this post is a panel from that report.

The candidate: a customer-support agent with six tools (`lookup_customer`, `lookup_order`,
`check_refund_policy`, `issue_refund`, `escalate_to_human`, `search_kb`), moving from
`gemini-3.7-pro` to `gemini-3.7-flash`. The suite: 120 examples recorded in production by the
EvalShift capture SDK, promoted into a golden JSONL suite with `evalshift capture sync`, and
sliced by tag according to what the conversation was about.

| Slice | Examples | What is in it |
| --- | --- | --- |
| `routine` | 42 | order status, shipping, account questions |
| `refund` | 26 | refund and return requests |
| `security` | 24 | account access, password and payment-method changes |
| `customer_lookup` | 16 | requests that need a customer record first |
| `text_only` | 12 | greetings, thanks, off-topic |

One `evalshift compare` command, real API calls on both sides, and the report opened on:

```report
{
  "kind": "verdict",
  "verdict": "pass",
  "summary": "10 of 10 budgets within policy.",
  "rates": { "equivalent": 90.1, "improved": 6.3, "regressed": 3.6 },
  "cards": [
    {
      "eyebrow": "Advisory signal",
      "value": "3.6%",
      "unit": "regression rate",
      "tone": "ok",
      "note": "Below the max_overall_regression_rate of 5%. 16 of 444 scored comparisons, none above medium severity."
    },
    {
      "eyebrow": "Economics",
      "value": "-71.2%",
      "unit": "cost",
      "tone": "ok",
      "note": "$1.9248 → $0.5544. Latency -61.9%. Both inside +0% cost / +30% latency."
    }
  ],
  "caption": "EvalShift's verdict card. Ten budgets, all held; the regression rate and the economics sit beside it so nobody has to scroll to find out what the pass cost."
}
```

```report
{
  "kind": "strip",
  "cells": [
    { "label": "Examples", "value": "120" },
    { "label": "Calls", "value": "240", "note": "0 cached" },
    { "label": "Failed / truncated", "value": "0 / 0", "tone": "ok" },
    { "label": "Total cost", "value": "$2.4792" },
    { "label": "Latency Δ", "value": "-61.9%", "tone": "ok" },
    { "label": "Avg score Δ", "value": "+0.011", "tone": "ok" }
  ],
  "caption": "The run strip. Same 120 examples on both sides, nothing failed or truncated, so every comparison below is over the full suite. EvalShift excludes truncated and errored calls from the statistics, so this row is worth checking first."
}
```

Cheaper by 71.2%, faster by 61.9%, verdict PASS. The economics look like the last post's. The
difference is everything that was decided before the run.

## The numbers were written before the run

EvalShift's `migration_policy` is a block in `evalshift.yaml` with seven budgets: overall
regression rate, critical regression count, equivalent-or-better rate, tool-argument drift,
tool-selection divergence, cost increase and latency increase. A `slices` map under it lets any
slice override any budget, inheriting the top-level value where it doesn't. Evaluators are
configured per suite; budgets are set once and tightened per slice.

`evalshift init --profile cost-reduction` scaffolds a starting policy: 2% overall regression
rate, zero critical regressions, 97% equivalence, 1% tool-argument drift, 5% cost increase, 30%
latency increase. That is a starting point, not a decision. This is what the project actually ran
with:

```yaml
migration_policy:
  max_overall_regression_rate: 0.05
  max_critical_regressions: 0
  min_equivalence_rate: 0.90
  max_tool_argument_drift: 0.10
  max_tool_divergence: 0.05
  max_cost_increase: 0.0
  max_latency_increase: 0.30
  slices:
    security:
      max_overall_regression_rate: 0.0
      max_tool_divergence: 0.0
    refund:
      max_tool_argument_drift: 0.0
```

Three things moved between the profile and this file.

### Check every rate against the suite size

A rate over *n* rows can only move in steps of 1/*n*. This suite has 108 tool-argument rows, so
one drifted row is 0.93%, and a 1% budget is a zero-tolerance budget wearing a percentage.
EvalShift prints a recommendation when a budget is below the granularity of its denominator,
naming the budget, the value and the row count, and this one would have triggered it.

So the rule I use: if I mean zero, I write `0.0`. Where I don't, I set a number the sample can
resolve. 10% argument drift on 108 rows is ten reworded search queries, which is what drift on
this agent looks like in practice.

### Tolerance where wording lives, zero where money and access live

The overall regression rate went *up*, from 2% to 5%. The pairwise LLM judge is blocking on this
project, and a judge flips on rewording; 2% of 444 comparisons is nine judge calls deciding the
migration. 5% leaves room for the judge to disagree about phrasing without anything real being
allowed through.

The strictness moved into the slices instead. `security` gets zero regressions and zero
tool-selection divergence: a model that starts routing an account-access request to a different
tool does not get a percentage. `refund` gets zero argument drift: an order id or an amount that
drifts is a wrong refund, not a reworded one.

The slices that are *not* in the policy matter too. `customer_lookup` has 16 examples. EvalShift
tests a comparison below 20 pairs but flags it uncertain, so that slice inherits the top-level
budgets and gets no tighter ones. It is the slice I would grow before I tightened it.

### The cost budget is the reason for the migration

`max_cost_increase: 0.0`. The point of the exercise is to spend less. A candidate that costs more
has failed before any quality number is read, and a policy should say so instead of leaving it to
whoever reads the economics card. EvalShift measures cost and latency from the run's own calls,
so these two budgets gate even when no quality evaluator does.

Two evaluator decisions sit alongside the budgets. Every EvalShift evaluator carries a
`blocking` flag; advisory (`blocking: false`) results are reported and ranked but never change
the verdict. The judge is `blocking: true` here; `init` writes `false` because at a dozen
examples judge noise would decide the verdict, and at 120 with an audited criterion it earns its
vote. `semantic` stays advisory. It measures wording, and wording is the one thing this
migration was allowed to change.

## What the run measured

```report
{
  "kind": "budgets",
  "rows": [
    { "name": "Overall regression rate", "id": "max_overall_regression_rate", "scope": "overall", "observed": "3.6%", "limit": "≤ 5.0%", "tone": "ok", "note": "16 of 444 · 95% CI 2.2–5.8%" },
    { "name": "Critical regressions", "id": "max_critical_regressions", "scope": "overall", "observed": "0", "limit": "≤ 0", "tone": "ok", "note": "of 444" },
    { "name": "Equivalent-or-better rate", "id": "min_equivalence_rate", "scope": "overall", "observed": "96.4%", "limit": "≥ 90.0%", "tone": "ok", "note": "428 of 444" },
    { "name": "Tool-argument drift", "id": "max_tool_argument_drift", "scope": "overall", "observed": "4.6%", "limit": "≤ 10.0%", "tone": "ok", "note": "5 of 108 tool-argument rows" },
    { "name": "Tool-selection divergence", "id": "max_tool_divergence", "scope": "overall", "observed": "2.8%", "limit": "≤ 5.0%", "tone": "ok", "note": "3 of 108 divergence rows" },
    { "name": "Cost increase", "id": "max_cost_increase", "scope": "overall", "observed": "0.0%", "limit": "≤ 0.0%", "tone": "ok", "note": "cost fell 71.2%" },
    { "name": "Latency increase", "id": "max_latency_increase", "scope": "overall", "observed": "0.0%", "limit": "≤ 30.0%", "tone": "ok", "note": "latency fell 61.9%" },
    { "name": "Overall regression rate", "id": "max_overall_regression_rate", "scope": "security", "observed": "0.0%", "limit": "≤ 0.0%", "tone": "ok", "note": "0 of 96" },
    { "name": "Tool-selection divergence", "id": "max_tool_divergence", "scope": "security", "observed": "0.0%", "limit": "≤ 0.0%", "tone": "ok", "note": "0 of 24" },
    { "name": "Tool-argument drift", "id": "max_tool_argument_drift", "scope": "refund", "observed": "0.0%", "limit": "≤ 0.0%", "tone": "ok", "note": "0 of 26" }
  ],
  "caption": "Every budget against its limit, as EvalShift reports them. Seven at the top level, three on the two slices where a regression is a wrong action rather than a reworded one."
}
```

The first row is the one worth reading twice. 16 of 444 is 3.6%, and its 95% Wilson interval
runs from 2.2% to 5.8%. The interval crosses the limit.

EvalShift's rule is asymmetric on purpose. A breach fails only when the interval confirms it; a
breach the interval cannot confirm returns `inconclusive`, because the suite was too small to
say. A budget the observation held is conclusive however wide its interval, because a wide
interval must never downgrade a clean run. This budget held, so it passes, and the interval is
printed so the reader knows how much room there was.

The three slice budgets held at exactly zero. Zero on 96 comparisons means something. Zero on 12
would not, which is why `text_only` has no slice budget at all.

## By evaluator

Four evaluators scored this run. EvalShift's `tool_selection` evaluator reads the recorded
traces and scores two axes: conformance, where each side is graded against the suite's recorded
tool calls, and divergence, where the target is graded against what the source did.
`tool_arguments` scores argument values field by field against the expected call. The
`llm_judge` is pairwise and sees the two outputs as anonymous A and B. `semantic` is embedding
similarity between the two outputs. The first three are blocking; the last is advisory.

```report
{
  "kind": "evaluators",
  "rows": [
    { "name": "Routing — conformance", "id": "routing · tool_selection.conformance", "axis": "each side graded against the suite's recorded tool calls", "n": 108, "delta": "+0.046", "effect": "0.28", "magnitude": "small", "ci": "[0.09, 0.47]", "confidence": "likely", "severity": "improved" },
    { "name": "Routing — divergence", "id": "routing · tool_selection.divergence", "axis": "the target graded against what the source did", "n": 108, "delta": "-0.028", "effect": "0.17", "magnitude": "negligible", "ci": "[-0.36, 0.02]", "confidence": "unclear", "severity": "none" },
    { "name": "Routing args", "id": "routing_args", "n": 108, "delta": "-0.004", "effect": "0.03", "magnitude": "negligible", "ci": "[-0.22, 0.16]", "confidence": "unclear", "severity": "none" },
    { "name": "LLM judge: equivalence", "id": "llm_judge.equivalence", "n": 120, "delta": "+0.029", "effect": "0.12", "magnitude": "negligible", "ci": "[-0.06, 0.30]", "confidence": "unclear", "severity": "none" },
    { "name": "Semantic similarity", "id": "semantic", "advisory": true, "n": 120, "delta": "-0.041", "effect": "0.44", "magnitude": "small", "ci": "[-0.62, -0.26]", "confidence": "likely", "severity": "medium", "blurb": "Reported, not gating: blocking is false." }
  ],
  "caption": "Overall, by evaluator. Each row is a paired test over that evaluator's deltas: effect size with a 95% interval, and a confidence label from the Benjamini-Hochberg corrected p-value. Four blocking rows say equivalent or improved; the one regression is on the advisory evaluator that measures wording."
}
```

Four blocking rows say equivalent or improved. The one that regressed is advisory, and it
measures the one thing this migration was allowed to change. Had `semantic` been blocking, the
same run would have come back `conditional_pass` on a medium-severity regression in phrasing.
That is not a reason to make every evaluator advisory. It is a reason to decide, per evaluator,
whether what it measures is something you are willing to block a migration on.

The conformance row says the candidate matched the suite's recorded tool calls *more often* than
the model that produced them: 11 examples improved, one regressed. Most of the eleven were refund
requests where the source went straight to `issue_refund`.

## The diffs I still read

A pass is not permission to skip the diff. For every flagged example, the EvalShift report shows
the reason it was flagged, the tool calls on each side, the argument-level diff, and the
conversation context that led to it. Two examples from this run, one from each side of the
ledger.

```report
{
  "kind": "example",
  "id": "cap_6b1e40f2a9c34d0b8e7d2a5f31c9e804",
  "turn": 1,
  "what": "Routing — conformance",
  "delta": "+0.500",
  "tone": "ok",
  "why": {
    "scores": "source 0.50 → target 1.00 (0–1)",
    "label": "IMPROVED",
    "text": "The suite expected check_refund_policy before issue_refund. The source went straight to issue_refund; the target called both, in order."
  },
  "tools": { "source": "issue_refund", "target": "check_refund_policy, issue_refund" },
  "source": {
    "score": "0.500",
    "tone": "bad",
    "calls": [
      { "tool": "issue_refund", "args": "{\"order_id\": \"ord_58213\", \"amount\": 42.0, \"reason\": \"damaged on arrival\"}" }
    ],
    "final": "Done — I've refunded $42.00 to your original payment method. You'll see it in 3–5 business days."
  },
  "target": {
    "score": "1.000",
    "tone": "ok",
    "calls": [
      { "tool": "check_refund_policy", "args": "{\"order_id\": \"ord_58213\"}", "mark": "extra" },
      { "tool": "issue_refund", "args": "{\"order_id\": \"ord_58213\", \"amount\": 42.0, \"reason\": \"damaged on arrival\"}" }
    ],
    "final": "Your order is inside the 30-day window, so I've refunded $42.00 to your original payment method. Expect it in 3–5 business days."
  },
  "diffs": [
    "extra tool: Target added check_refund_policy at position 1.",
    "matched: issue_refund arguments are identical on both sides."
  ],
  "caption": "An improvement, as EvalShift's per-example diff shows it. The recorded ground truth expected the policy check first; the source model that produced the traffic skipped it and the candidate did not."
}
```

The refund went out either way, and the final text on both sides would pass any output check
you care to write. The trace is where the difference lives: one side verified before acting and
the other did not. This is the class of change a text evaluator cannot see in either direction,
and the reason EvalShift's tool-call evaluators score the trace rather than the prose.

```report
{
  "kind": "example",
  "id": "cap_9d27c5e1b04a4f6e9a3c7d8e2f105b6c",
  "turn": 0,
  "what": "Routing args",
  "delta": "-0.160",
  "tone": "bad",
  "why": {
    "scores": "source 1.00 → target 0.84 (0–1)",
    "label": "ARGUMENT_VALUE_DRIFT",
    "text": "search_kb.query scored 0.84 against the expected value under the semantic strategy, below the 0.9 drift floor. Same tool, same intent, different wording."
  },
  "tools": { "source": "search_kb", "target": "search_kb" },
  "source": {
    "score": "1.000",
    "tone": "ok",
    "calls": [
      { "tool": "search_kb", "args": "{\"query\": \"change shipping address after order placed\"}" }
    ],
    "final": "You can change the address until the order ships — open the order and tap “Edit address”. Once it's marked shipped, reply here and I'll reroute it with the carrier."
  },
  "target": {
    "score": "0.840",
    "tone": "bad",
    "calls": [
      { "tool": "search_kb", "args": "{\"query\": \"edit delivery address existing order\"}" }
    ],
    "final": "Until the order ships you can edit the address yourself from the order page. If it has already shipped, let me know and I'll set up a carrier reroute."
  },
  "diffs": [
    "argument drift: search_kb.query scored 0.84 against the expected query (drift floor 0.9).",
    "same tool set: no calls added or removed."
  ],
  "caption": "A regression the budget was built to tolerate. Five of 108 tool-argument rows drifted like this one; the limit was ten."
}
```

A reviewer would call those the same query. The evaluator scored it 0.84 against a 0.9 floor and
counted it as drift, five rows out of 108 did the same, and the budget said five is fine. That is
what the 10% was for. The same five rows in the `refund` slice would have failed the run, and
that was also decided in advance.

## What EvalShift did in this run

The whole run, as a list of the product's parts, in the order they were used:

- **The capture SDK** recorded the agent's real conversations in production, tool calls
  included, and `evalshift capture sync` promoted them into a frozen golden JSONL suite with
  the tool evaluators written from what the captures actually contained.
- **`evalshift compare`** replayed every example against the source and the target model, paired
  per example, with the same inputs, tools and context on both sides.
- **The tool-call evaluators** (`tool_selection`, `tool_arguments`) scored the traces, the
  **pairwise LLM judge** scored the outputs, and **`semantic`** measured drift in wording,
  advisory only.
- **Paired statistics** turned each evaluator's deltas into an effect size, a 95% interval and
  a corrected confidence label, so a two-point average drop and a real regression are told
  apart mechanically.
- **`migration_policy`** in `evalshift.yaml` held seven budgets, three of them tightened to zero
  on the `security` and `refund` slices, and every proportion budget was judged with a Wilson
  interval that can return `inconclusive` instead of a false fail.
- **`report.html`** was written locally, verdict first, with a per-example diff for everything
  flagged. Nothing left the machine.
- **`--policy-gate`** made the verdict an exit code, which is what the EvalShift GitHub Action
  uses to block a pull request on the same policy once the suite runs in CI.

## What a pass proves

Not that nothing changed. 16 comparisons regressed and 28 improved; the agent phrases search
queries differently and checks the refund policy more often than it used to.

It proves that every change stayed inside limits that were written down before anyone saw a
number, on a suite that was frozen before the run. That is the entire claim, and it is enough to
act on, because there is nothing left to negotiate: the argument about what counts as acceptable
happened in the YAML, not in the meeting after the report.

What happens next is the boring part, which is the point. The model string changes in production.
The suite does not. It runs again on the next pull request through the EvalShift GitHub Action,
gated on the same policy, against the new baseline.

```bash
evalshift compare --suite-name support_routing --to gemini-3.7-flash --policy-gate --open
```

If the previous post was the reason to run the comparison, this one is the reason to write the
policy first.]]></content:encoded>
    </item>
    <item>
      <title>The new model was 59% cheaper and 75% faster. I still wouldn&apos;t ship it.</title>
      <link>https://www.evalshift.dev/blog/cheaper-faster-and-still-a-fail</link>
      <guid isPermaLink="true">https://www.evalshift.dev/blog/cheaper-faster-and-still-a-fail</guid>
      <pubDate>Sat, 12 Sep 2026 00:00:00 GMT</pubDate>
      <description>A candidate model cut cost 59% and latency 75% with zero failed calls. The migration report still said FAIL, because the agent&apos;s behavior changed.</description>
      <content:encoded><![CDATA[The migration looked like an obvious win on economics. In one EvalShift run, the candidate model
was:

- 58.8% cheaper
- 75.2% faster
- 0 failed or truncated calls

![Run economics table: source and target each made 16 calls with no failures or truncations; the target cost $0.0741 against $0.1799 and averaged 773 ms against 3.1 s.](/blog-images/blog1/blog1-3.png "Run economics. Same 16 calls on both sides, no failures, and every number a finance reviewer cares about moved the right way.")

Then the migration report returned:

**FAIL**

Not because the model crashed. Not because the API changed. Not because the output stopped
parsing.

It changed what the agent actually did.

![Migration verdict card: FAIL, 5 of 7 budgets within policy, 61.4% equivalent, 25.0% improved, 13.6% regressed, cost −58.8%, latency −75.2%.](/blog-images/blog1/blog1-1.png "The verdict card. Five of seven budgets passed; the two that did not are the two that describe behavior.")

## What failed

I replayed 16 real examples against the source and candidate models. Across the run, 13.6% of
evaluated comparisons regressed.

The overall regression-rate budget actually passed.

Two other checks did not:

- Tool-selection divergence hit 25%, against a 10% limit.
- The overall equivalence rate fell to 61.4%, below the required 75%.

![Report verdict and findings: the target breached max_tool_divergence at 25.0% against a 10% ceiling and fell below min_equivalence_rate at 61.4% against 75%; findings list searching instead of answering, confirming actions it never performed, and routing to the wrong tool.](/blog-images/blog1/blog1-2.png "The written verdict and its findings. The recommendation is not to migrate until the tool-selection and semantic regressions are fixed.")

That distinction matters. Looking only at latency, cost, successful requests, or even a few
manually inspected outputs would have made this migration look pretty attractive.

The behavioral diff told a different story.

## The failures weren't cosmetic

One example asked for an opinion on working late.

The source model answered the question directly.

The candidate instead issued a `search_web` call and returned no text.

![Per-example diff for the working-late question: the source called no tools and answered in two sentences; the target called search_web with the query "is working late worth it productivity well-being" and produced no final text. Flagged as TOOL_SELECTION_DRIFT.](/blog-images/blog1/blog1-5.png "The diff for that example. Source trace on the left, target trace on the right, and the reason it was flagged at the top.")

Other cases had the same general problem in different forms:

- an action was reported as completed even though the corresponding tool was never called;
- a request was routed to a different tool;
- an unnecessary tool call appeared where the source answered directly.

These are not necessarily signs that the candidate model is "bad." They show that changing the
model changed the application.

For an agent, the model is part of the control flow.

## Aggregate quality can hide this

The evaluator breakdown made the tradeoff clearer.

Routing arguments and routing conformance were broadly equivalent in this run, while
tool-selection divergence, semantic similarity, and the pairwise equivalence judge showed
regressions.

![Per-evaluator table over 16 examples: routing divergence regressed (score delta −0.167, likely), routing args equivalent (+0.179, unclear), routing conformance equivalent (+0.047, unclear), the LLM equivalence judge regressed critically (−0.438, certain), and semantic similarity regressed (−0.214, likely).](/blog-images/blog1/blog1-4.png "Overall, by evaluator. Two rows say equivalent, three say regressed, and each comes with its own effect size, confidence interval and confidence label.")

A single average score would flatten all of that into one number.

For a migration, I care more about the question:

*What changed, on which examples, and is that change acceptable for this application?*

## This is why I built EvalShift

I kept running into model migrations that were tested roughly like this:

change model → try several prompts → outputs look fine → ship.

That works until the difference is something subtle like an extra tool call, a missing action, or
different routing behavior.

EvalShift replays the same suite against the source and target model, compares the behavior, and
produces the migration decision and individual diffs before the model string gets changed in
production. The CLI is local-first and produces a self-contained HTML report; hosted upload is
optional.

The interesting result from this run wasn't that the candidate was worse.

It was that "59% cheaper and 75% faster" wasn't enough information to decide whether the migration
was safe.]]></content:encoded>
    </item>
    <item>
      <title>What actually breaks when you switch LLMs</title>
      <link>https://www.evalshift.dev/blog/what-breaks-when-you-switch-llms</link>
      <guid isPermaLink="true">https://www.evalshift.dev/blog/what-breaks-when-you-switch-llms</guid>
      <pubDate>Thu, 20 Aug 2026 00:00:00 GMT</pubDate>
      <description>A model swap is a behavior change with no diff to review. What moves besides the answer text, and why the worst regressions read as correct.</description>
      <content:encoded><![CDATA[Changing the model behind an AI feature looks like the smallest change you will make all week. One
string moves from `gpt-x` to `gemini-y`. The new model is cheaper, or faster, or ahead on the
benchmark someone linked in Slack. You try a handful of prompts, the answers read fine, you ship.

The reason this keeps going wrong is that a model swap is not a config change. It is a behavior
change, delivered through a config file, with no diff for anyone to review.

## The change surface is bigger than the answer text

Swapping the model can move any of these independently, and most teams only look at the last one:

| What moves | How it usually surfaces |
| --- | --- |
| which tools the agent calls | a step silently stops happening |
| the order of those calls | a check runs after the action it was meant to gate |
| tool arguments | right tool, wrong amount, wrong id, wrong units |
| structured output | your parser throws, or worse, doesn't |
| refusal behavior | the model declines work it used to do |
| verbosity and format | downstream regex and UI truncation start missing |
| latency | p95 doubles, nobody attributes it to the swap |
| tokens and cost | the cheaper model turns out to be the pricier one per task |
| answer quality | the only one the playground actually shows you |

A migration can improve one row and wreck another. A model that answers just as well but issues one
extra tool call per turn is not a cost reduction. You will not learn that by reading answers.

## The worst regressions are invisible in the output

Take an agent that is supposed to do this:

```text
lookup_order("A-339")
issue_refund("A-339", 29.99)
```

The new model does this instead:

```text
lookup_order("A-339")
```

and replies:

> Your refund has been processed.

Every text-based check passes. The sentence is fluent, on topic, and exactly what the old model said.
The refund did not happen. If you are scoring outputs, this regression is not merely hard to catch —
it is invisible by construction, because the output is correct and the behavior is not.

The same shape covers most of the expensive failures: the `verify_payment` call that stops firing,
the retry loop that starts, the confirmation step that moves after the write. What changed is the
trace. The prose stayed still.

## You are usually changing two things at once

Prompts are coupled to models. The system prompt in production has been tuned — often over months,
often by accretion — against one model's quirks. Point it at a different model and some of that
tuning becomes dead weight and some becomes actively harmful.

So the honest migration usually involves editing the prompt too, and now the comparison has two
independent variables in it. When quality moves, nobody can say whether the model did it or the
rewrite did.

The fix is boring: change one thing per run. Baseline old model with old prompt. Run new model with
old prompt — that is the model's effect, unflattering as it may be. Then tune the prompt for the new
model and run again, against the same frozen cases. Two comparisons, each interpretable, instead of
one that isn't.

## Hand-written test cases test the paths you already handle

Writing eval prompts by hand feels productive and produces a suite shaped like your mental model of
the product. That is the problem. The prompts you can think of are the interactions you already
understand well enough to have handled.

Real failures come from the inputs you would never have written down: the eleven-turn conversation
where a constraint set in turn three quietly expires, the message that arrives with half the context
missing, the turn right after a tool returned an error, the customer typing in a language your
template never anticipated. You cannot reconstruct those from memory. You have to record them.

Which gives a workflow, independent of what you use to run it:

```text
real agent behavior
        ↓
capture representative cases
        ↓
freeze a golden suite
        ↓
baseline model vs candidate model, same inputs
        ↓
compare behavior, not just text
        ↓
ship or reject
```

The freezing step is the one people skip. If cases are still being edited while models are being
compared, two things are moving and the diff between them describes neither.

## An average is not a verdict

Model A scores 0.91, model B scores 0.89, and someone screenshots it into the migration thread. That
number cannot carry the decision. With a couple of dozen noisy samples, a two-point gap is well
inside what you would see running the *same* model twice.

Two things make it a real comparison. Run paired — every case against both models, then subtract per
case, so the fact that some cases are inherently harder cancels instead of drowning the signal. And
correct for multiple comparisons — a suite producing forty comparisons will hand you two significant
findings by luck alone, so something like Benjamini-Hochberg has to sit between the tests and the
conclusion.

But statistical significance is not product importance, and this is where the reasoning usually stops
one step early. A semantic-similarity drop of 0.02 can be real, reproducible, significant, and
irrelevant. One missing `verify_payment` call across two hundred cases is statistically nothing and
operationally a serious problem. Significance tells you the effect exists. It has no opinion about
whether you should care.

## Decide what matters before you see the numbers

Which is why the last artifact of a migration is a written policy — thresholds agreed while the
result is still unknown, so the verdict is read off rather than negotiated:

```text
missing critical tool        -> fail
invalid JSON above threshold -> fail
latency +10%                 -> warn
small semantic delta         -> ignore
```

Write that after the run and the thresholds bend around the number you were hoping for. Everyone does
this; nobody means to.

A policy also gives you a fourth verdict worth having explicitly: *inconclusive*. Not "pass",
not "fail" — "this suite is too small to tell you". Teams that collapse that into a pass ship
regressions they had the evidence to catch, one underpowered comparison at a time.

## Disclosure, and the part that survives it

I build [EvalShift](/), which does exactly this: capture real agent runs, freeze them into a golden
suite, run both models paired, score tool calls and arguments and structure alongside output quality,
and gate the pull request on a policy you wrote in advance. So take the tool mention as interested.

The argument underneath it is not:

If an LLM decides behavior in your system, then changing the model is a behavior change, and it
deserves the same treatment as a dependency bump that alters runtime semantics — frozen test cases,
a before-and-after run, and a decision rule written while you still have no stake in the answer.

Not a string edit in a config file, followed by hope.

## Keep reading

- [How to test an LLM model migration before you ship it](/blog/test-llm-model-migration-before-you-ship) —
  the same argument as a method, step by step.
- [Evaluating agent tool calls](/blog/evaluating-agent-tool-calls) — what text evaluators structurally
  cannot see, and the four evaluators that can.
- [Build a golden eval suite from production traffic](/blog/build-a-golden-suite-from-production-traffic) —
  turning recorded runs into the frozen suite this all depends on.]]></content:encoded>
    </item>
    <item>
      <title>Prompt edits deserve the same gate as a model swap</title>
      <link>https://www.evalshift.dev/blog/prompt-regression-testing</link>
      <guid isPermaLink="true">https://www.evalshift.dev/blog/prompt-regression-testing</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>A prompt edit changes behavior the way a model swap does. Where the prompt text lives decides whether CI diffs it or passes on silence.</description>
      <content:encoded><![CDATA[A model swap gets a meeting. A prompt edit gets a commit. Both change what your product says to
users, and only one of them is reviewed like it matters — usually because a prompt diff is one
paragraph of English that reads fine and proves nothing.

The gate you already built for model migrations works on prompt edits unchanged: a golden suite, a
paired run, a diff against the base branch. But whether a prompt edit actually reaches that gate
depends on something most teams never look at — where the prompt text lives.

## The gate, briefly

On a pull request the EvalShift action runs the suite, uploads the run, asks the Cloud API for a
compatible baseline run on the base branch, fetches the server-side diff, updates one PR comment,
and sets the `evalshift/regression` commit status. Under `fail-on: regression` the check goes red
when the diff shows regressions.

The load-bearing word is *compatible*. The lookup scans recent `available` runs on the base branch
for the same suite, and skips every one whose `eval_config_hash` differs from the candidate's.
No compatible baseline means no diff — and no diff means the check passes.

## Where the prompt lives decides whether it is gated

`eval_config_hash` is a SHA-256 over the canonical JSON of your config snapshot: `version`,
`prompts`, `defaults`, `evaluators`, `slices`. The `prompts` block is in the hash. What is inside
that block depends on the detection mode:

```yaml
# detection: manual — the prompt body IS the config
prompts:
  - id: replay
    detection: manual
    content: "You are a support agent. Answer in at most three sentences.\n\n{input}"
    variables: [input]

# detection: python_string — the config only points at the body
prompts:
  - id: customer_routing
    detection: python_string
    path: prompts.py
    variable: AGENT_SYSTEM_PROMPT
    variables: [query]
```

With `detection: manual`, editing one word of the prompt changes `prompts[].content`, which changes
the config snapshot, which changes `eval_config_hash`. Your PR's run is now incomparable with every
run on `main`. The action finds no compatible baseline, reports exactly that in the comment, and
exits green.

With `detection: python_string`, `content` must be null — the config carries only `path` and
`variable`, and the body is AST-extracted from the Python file at run time. Edit the prompt and the
hash does not move: the baseline lookup matches, the server diffs the runs directly, and the
regression shows up as a regression.

> A permanently green EvalShift check usually means no comparable baseline exists, not that the
> suite is passing. If the check went green the same week someone rewrote the system prompt inline,
> that is the mechanism, not a coincidence.

The same rule governs everything else in the hash: retuning `defaults`, adding an evaluator, or
renaming a slice all re-baseline you on purpose. Prompt wording is the one that gets edited weekly
and the one nobody expects to sever the comparison.

## What still changes when the text changes

Keeping the prompt out of the config hash does not make edits invisible to the pipeline:

- The response cache is keyed on `{model, prompt, inputs, temperature, max_tokens[, history]}`, so
  an edited prompt is a cache miss. You cannot accidentally score new wording with old outputs.
- Every example is validated against every prompt — template variables covered — *before* any model
  call is dispatched. Dropping `{order_id}` from the template fails the run in seconds rather than
  after $9 of calls.
- `python_string` extraction never imports or executes your code. It AST-walks for a module-level
  string assignment and takes the last one; f-strings, concatenation, `.format()`, and function
  calls are rejected rather than evaluated.

That last rule has a sharp edge. A prompt assembled at runtime cannot be extracted, and the
documented workaround is to paste it into the config as `detection: manual` — which puts the body
back inside the hash and back into re-baselining on every edit. Prefer refactoring the template so
the literal is the literal and the variables are `{placeholders}` the suite fills in.

## Wire it once

Keep one copy of the prompt, in the application code that ships it, and point the config at it:

```yaml
prompts:
  - id: customer_routing
    detection: python_string
    path: app/prompts.py
    variable: AGENT_SYSTEM_PROMPT
    variables: [query]
```

Your app imports `AGENT_SYSTEM_PROMPT`; EvalShift reads the same file. There is no second copy to
drift, and a reviewer looking at the PR diff sees the prompt change and the eval result side by
side. Tool schemas follow the same no-second-copy rule by a different route: they live on the
suite examples, not in the config — `evalshift capture sync` records each example's toolset as a
content-addressed `toolset_ref` sidecar — so they never enter `eval_config_hash` either.

Then let the run answer the question the prose cannot. Both sides score paired per example, so
"the new prompt is more concise" becomes a delta per evaluator with a severity attached, and the
migration policy turns it into `pass`, `conditional_pass`, `fail`, or `inconclusive` without a
meeting. Roll out with `fail-on: never` for a week to collect baselines, then switch to
`fail-on: regression` once the comments match your judgement.

## The review question worth adopting

For every PR that touches a prompt: *did this run diff against a baseline, or against nothing?*
The comment answers it in one line. A prompt edit that produced no comparison has not been
tested — it has only been observed to compile.

## Keep reading

- [LLM regression testing in CI](/blog/llm-regression-testing-in-ci) — the gate itself, end to end.
- [How many eval cases do you need?](/blog/how-many-eval-cases-do-you-need) — sizing the suite that
  gate reads.
- [Evaluating agent tool calls](/blog/evaluating-agent-tool-calls) — what changes in the trace when
  a system prompt changes.
- [Configuration](/docs/configuration) and [Baselines](/docs/baselines) — the config fields and the
  baseline model in full.]]></content:encoded>
    </item>
    <item>
      <title>How many eval cases do you need?</title>
      <link>https://www.evalshift.dev/blog/how-many-eval-cases-do-you-need</link>
      <guid isPermaLink="true">https://www.evalshift.dev/blog/how-many-eval-cases-do-you-need</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <description>Suite size is the wrong question: n is counted per prompt, evaluator and slice. What the analysis drops, and how many pairs a verdict needs.</description>
      <content:encoded><![CDATA[Everyone asks the question the same way: how many examples does a golden suite need? Fifty?
Two hundred? The number that actually decides whether your run can say anything is not the size of
the suite. It is `n` inside one comparison — and the suite is chopped into comparisons before a
single statistic is computed.

This is where most first suites disappoint. Two hundred examples feel serious, come back
`inconclusive`, and nobody can explain why. Here is how `n` is counted, what is dropped before
counting, and how to size a suite so the verdict is earned rather than lucky.

## `n` is per comparison, not per suite

The analysis groups paired deltas — `target_score - source_score`, per example — by the triple
(`prompt_id`, `evaluator_name`, `slice_name`). Every one of those groups is tested on its own, and
every one has its own `n`.

So a 200-example suite run across 3 prompts, scored by 4 evaluators, sliced 3 ways is not one
comparison with n=200. It is up to 36 comparisons, each drawing on the rows that match its slice.
Give the smallest slice 15 examples and that column of the matrix is stuck at n=15 no matter how
impressive the suite total looks.

Three constants decide what happens next:

- `MIN_N_FOR_TEST = 5` — below five paired observations the comparison is skipped entirely, with
  severity `insufficient`. It is not scored as "no change"; it is scored as "not measured".
- `MIN_N_RELIABLE = 20` — between 5 and 20 the comparison is tested but flagged uncertain.
- Zero variance (std < 1e-9) — skipped with severity `none`, because every delta was identical.

Twenty per group is the line where a result stops carrying an asterisk. Work backwards from the
groups you care about, not forwards from a round suite size.

## Rows evaporate before `n` is counted

The `n` a comparison is tested at is smaller than the number of examples you wrote, and the gap is
deliberate:

- Rows an evaluator measured nothing on — a tool-only turn handed to a text evaluator, say — write
  no score at all, so they are absent from `n` before it is computed, and noted as "K of N rows
  not applicable". They also leave the slice aggregates and the policy metrics.
- Truncated calls — the ones that hit `max_tokens` — are excluded from statistics, because a
  cut-off answer scores badly for a reason that has nothing to do with the model's quality.
- Evaluator-side failures (a judge call that broke, an embedding request that errored) are stored
  as `errored` and excluded. A flaky judge shrinks `n` and drifts a comparison toward
  `insufficient`, which is visible, instead of poisoning the mean, which is not.

> When nothing survives, the comparison reports n=0 with a note prefixed `nothing measured:` —
> never severity `none`. Unmeasured is not equivalent, and a blocking evaluator that measured
> nothing never enforced its gate. That is exactly why the policy downgrades an otherwise-passing
> verdict to `conditional_pass` when a gating comparison carries that note.

## Slices are not free

Slicing is how you find the regression that only hits Spanish, or only hits tool-heavy turns. It
costs twice.

First, it splits the same rows into more, smaller groups — the fastest way to convert a healthy
comparison into three uncertain ones. Second, every testable comparison goes through a
Benjamini-Hochberg FDR correction at α=0.05 across **all** comparisons in the run. More
comparisons means each one clears a stricter bar to stay significant.

Slice because a subgroup can move independently and you would act on it, not because the tag
existed in your data. (Slices holding identical (prompt, evaluator, example) triples collapse to
one, so duplicating a slice under two names buys nothing.)

## Effect size, not just p, sets severity

A statistically significant result is not automatically a blocking one. Severity comes from the
FDR-corrected p-value together with paired Cohen's d — `mean(deltas) / std(deltas, ddof=1)`, with
a 95% CI that is analytical after a t-test and a seeded percentile bootstrap after Wilcoxon:

- `critical` — a regression with p < .01 and |d| > 0.8
- `high` — significant, |d| > 0.5
- `medium` — significant, |d| > 0.2
- `low` — significant, small effect
- `improved` — significant and positive

That ladder is the practical sizing question restated: a suite sized to catch only |d| > 0.8 will
sail past the medium drift that annoys users every day. Small effects need more pairs, and no
amount of confidence in the config substitutes for them.

Which test runs is decided for you: Shapiro-Wilk on the deltas at α=0.05 picks a paired t-test when
they look normal and a Wilcoxon signed-rank test when they do not — skipped above n=5000, where the
CLT justifies a t-test outright. Judge deltas confined to +1/0/-1 usually fail that screen and land
on Wilcoxon.

## Rate budgets need a wider suite than tests do

`migration_policy` rate budgets — `max_overall_regression_rate`, `min_equivalence_rate` — are
Wilson-confidence-interval aware at 95%. A breach fails the run only when the interval confirms it.
Breach with the interval still spanning the budget returns `inconclusive`, and the reason is the
suite, not the model.

Put concretely: a 3% regression-rate budget cannot be confirmed breached by a 30-example suite.
One bad example is 3.3%, and the interval around it is enormous. Cost and latency budgets are exact
and always conclusive, because they are ratios over measured calls rather than rates over
sampled outcomes.

## A sizing rule that survives contact

1. List the comparisons you would actually act on — the (prompt, evaluator, slice) triples where a
   regression would change your decision. That count, not the example count, is the shape of the
   run.
2. Target 20 paired observations in each of them. Below 20 you get an answer with a flag on it;
   below 5 you get no answer at all.
3. Add headroom for what gets dropped — non-applicable rows, truncations, evaluator errors. A 25%
   cushion is not paranoid on agent suites.
4. Start with fewer slices than you think you want. You can always split a slice later; you cannot
   un-spend the FDR budget on slices nobody read.
5. Check the bill before the run. A paired run is prompts × examples × 2 models, and the CLI asks
   for confirmation above a $10 estimate (`--yes` or `EVALSHIFT_NONINTERACTIVE` skips it). Iterate
   on suite *structure* with `evalshift validate` — it loads config, suite, and prompts and
   cross-checks them without a single model call — and let the 7-day response cache absorb the
   reruns where structure did not change.

The honest version of "how many cases do I need" is: enough that the comparisons you would act on
each hold twenty pairs after attrition. For most teams that is a smaller, sharper suite than the
one they were planning — and one that comes back with a verdict instead of a shrug.

## Keep reading

- [How to test an LLM model migration before you ship it](/blog/test-llm-model-migration-before-you-ship)
  — the paired run these statistics summarize.
- [Build a golden eval suite from production traffic](/blog/build-a-golden-suite-from-production-traffic)
  — where the examples come from in the first place.
- [When to trust an LLM judge](/blog/when-to-trust-an-llm-judge) — the evaluator most likely to
  shrink your `n` without telling you.
- [Methodology](/docs/methodology) — the statistics contract in full.]]></content:encoded>
    </item>
    <item>
      <title>Evaluating agent tool calls: what text evals can&apos;t see</title>
      <link>https://www.evalshift.dev/blog/evaluating-agent-tool-calls</link>
      <guid isPermaLink="true">https://www.evalshift.dev/blog/evaluating-agent-tool-calls</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>Agent behavior drifts in the trace, not the prose. The four tool-call evaluators, the modes worth changing, and the expectations not worth pinning.</description>
      <content:encoded><![CDATA[The output looked identical. That is the sentence at the center of most agent migration incidents.
Both models answered "I've refunded your order, you'll see it in three to five business days" — and
one of them called `issue_refund` before `lookup_order`, on an order id it had not verified. Text
evaluators cannot see that, because the thing that changed was never in the text.

Agent behavior is a trace: which tools, in what order, with what arguments, and how many times. It
drifts independently of prose quality, which is why a migration can pass every judge criterion you
wrote and still be the wrong deploy. This post is what to score on the trace, and how the four
tool-call evaluators divide that work.

## Why text evaluators go quiet exactly here

On a turn where both models answered with tool calls and no prose, there is nothing to embed and
nothing to judge. The `semantic` evaluator writes no record at all — rather than erroring on an
empty embedding input or inventing a score — and `llm_judge` does the same without spending a
judge call, since comparing two empty strings only ever returns a meaningless tie, and a
fabricated tie is indistinguishable from a judged one.

Both behaviors are correct, and together they mean your text evaluators contribute nothing on the
most agent-shaped turns in the suite. If tool calls are how your product does its work, the tool-call
evaluators are not an addition to your eval config. They are the config.

(The asymmetric case is still scored: a target that went silent where the source answered in prose is
exactly the regression `semantic` exists to catch.)

## The toolset travels with the example

Nothing on the prompt marks it as an agent prompt. The toolset rides on each golden-suite
*example* instead — a content-addressed `toolset_ref` pointing at a sidecar file under
`.evalshift/toolsets/`, or an inline `tools` list — so one suite freely mixes agent and text-only
rows under the same prompt. For any example whose toolset is non-empty, the orchestrator sends the
provider those tool definitions and records each response as a provider-agnostic `ToolTrace` —
ordered `ToolCall`s carrying `tool_name`, `arguments`, `call_id`, `parent_call_id`, and
`sequence_index`, plus `final_text` and refusal info.

A toolset accepts both provider shapes, as a flat list or `{"tools": [...]}`:

```yaml
- name: issue_refund
  description: Issue a refund on an existing order.
  input_schema:
    type: object
    properties:
      order_id:   {type: string}
      amount_usd: {type: number}
    required: [order_id, amount_usd]
```

The model client serializes to whatever the target provider expects, so one file serves Anthropic,
OpenAI, and Gemini alike — which matters, because a cross-provider migration is the case where a
hand-maintained second copy of your tool schemas drifts first.

A capture-first suite never has that second copy at all. `evalshift capture sync` writes one
sidecar per distinct toolset your captures recorded — content-addressed, so two captures offering
the same tools share one file — and stamps each example's `toolset_ref` to match: the schemas your
suite replays are the schemas your agent actually offered, recorded at the moment it offered them.
Hand-authored suites inline the same shape as the example's `tools:` list instead, and `tools: []`
is a real value — "this example's agent had no tools available" is a first-class assertion, not an
absence.

## The four evaluators

All of them are pure computation over recorded traces: no API calls, no cost, no opinion that can be
argued with. Configure all four and you are still paying less than one judge criterion.

### `tool_selection` — did it call the right tools?

Two independent axes, each writing its own record, because a migration asks two different
questions and the answers differ:

| Axis | Compares | Strategies |
| --- | --- | --- |
| `conformance` | Each side, absolutely, against the example's `expected_tools` | `expected` (default — matched **in order**), `expected_set` (order-insensitive multiset recall, for parallel fan-outs), `off` |
| `divergence` | The target against the source | `set` (default — Jaccard over the tool-name sets), `exact` (sequence equality), `first` (first call only), `off` |

Conformance is the axis that measures correctness: the source model is a baseline, not an oracle,
so both sides are graded against recorded ground truth — and a row where *both* sides missed it is
tagged `TOOL_GROUND_TRUTH_MISS` and left out of the policy rates, because ground truth the source
model itself fails is a broken harness, not a migration finding. Divergence measures drift, with
the source as its own 1.0 baseline; its default is `set` rather than `exact` so reordered identical
calls do not read as change. Turning both axes off is a config error.

Examples marked `expected_no_tools` score 1.0 on conformance if and only if zero calls were made.
That is the evaluator for "answer the policy question, don't hit the database."

One extra knob deserves to be used more than it is: `severity_floor: low|medium|high|critical` means
a regression on this evaluator can never be classified below the floor regardless of effect size. The
canonical agent migration failure — the candidate quietly stops calling `notify_security_team` on
security-sensitive tickets — is a small effect on a small slice, and a floor is what keeps it from
being filed as `low` next to a formatting nit.

### `tool_arguments` — same tools, different values?

Calls are matched greedily by `(tool_name, nearest sequence_index)`, then each argument field is
scored by a per-field strategy: `exact`, `subset`, `numeric` (relative error decaying linearly to
zero at `numeric_tolerance`, default `0.05`), or `semantic` (embedding cosine, borrowing the
configured `semantic` evaluator's model and cache — without one it degrades to `exact`). Fields you
do not list default to `exact`.

Two defaults carry real judgment:

- A field present on one side only scores **0.5**, not 0. Omitting an optional parameter is a real
  difference, not a wrong value. `optional_fields_scored: strict` restores the harsher 0.0.
- `against: expected` switches the whole comparison from drift-vs-source to correctness, scoring
  **both** sides against `expected_tools[].arguments`. Each expectation's `match_strategy` picks which
  keys get compared: `exact` compares the union, `subset` and `contains_per_field` compare only the
  recorded keys. An expected call the model never made scores 0; an example with no expected
  arguments is skipped neutrally at 1.0/1.0.

Default drift mode pins `source_score` at 1.0 by construction. That is fine for "did anything move"
and wrong for "is it right" — a hallucinating source scores a perfect 1.0 forever. If your suite came
from real captures, you have the ground truth; use `against: expected`.

### `tool_trace_structure` — the shape of the loop

Call-count drift within `call_count_tolerance` (default 1), parallelism match, `expected_tool_count`
when set, and refusal alignment. Each check toggles off independently via `check_call_count`,
`check_parallelism`, `check_refusals`.

Refusals are the sharp edge: a refusal mismatch forces severity to at least `high` and flags
`REFUSAL_REGRESSION`. A model that started refusing work it used to do, or stopped refusing work it
used to decline, is never a low-severity finding — and it is invisible to every evaluator that scores
output text, because a refusal is fluent, well-formed prose.

### `agent_trace` — for agents that run outside EvalShift

If your agent loop lives in LangChain, a custom orchestrator, or another language entirely, run the
model-call stage and then attach full timelines:

```bash
evalshift traces import <run-id> --source source_traces.jsonl --target target_traces.jsonl
```

Each line is one trace for a `(prompt_id, example_id, role)`, with the same event schema the capture
SDK writes. Then `agent_trace` scores order similarity (LCS-normalised), per-field argument equality
on matched calls, and — the check with no equivalent anywhere else — missing verification:

```yaml
evaluators:
  agent_trace:
    - name: safety
      check_missing_verification: true
      verification_tools: [confirm_with_user]
      dangerous_tools: [delete_record, transfer_funds]
```

An extra dangerous call on the target flags `DANGEROUS_ACTION_DRIFT`; a dangerous call with no
preceding verification tool flags `MISSING_VERIFICATION_STEP`. Those two are worth encoding before
you need them, because they are the failures that turn a quality regression into an incident report.

## Read the categories, not the average

Regressions carry machine-readable labels that the report and the Cloud diff group by:
`TOOL_SELECTION_DRIFT`, `ARGUMENT_VALUE_DRIFT`, `TOOL_TRACE_STRUCTURE_DRIFT`, `TOOL_ORDER_DRIFT`,
`DANGEROUS_ACTION_DRIFT`, `MISSING_VERIFICATION_STEP`, `UNNECESSARY_TOOL_CALL`,
`REFUSAL_REGRESSION`, alongside the text-side ones.

The grouping is the point. "Agent quality dropped 4%" is not actionable and not even really a claim.
"Twelve `ARGUMENT_VALUE_DRIFT` on `issue_refund.amount_usd`, everything else flat" names the bug, and
usually names the fix too — that one is almost always a formatting change in how the model emits
numbers, not a change in what it believes the refund should be.

## Budget it in the policy

Two `migration_policy` fields exist specifically for the argument surface:

```yaml
migration_policy:
  max_tool_argument_drift: 0.01
  tool_argument_drift_floor: 0.9
```

The floor is what keeps the budget honest. Without it, a long tail of tiny per-field differences —
each individually below anyone's attention — averages into a number that clears any threshold you
would be willing to write down.

## What does not belong in the trace expectations

The failure mode on the other side is a suite so strict nobody can merge anything. Some of what a
capture records is behavior; the rest is circumstance.

- Timestamps, request ids, session tokens, and anything else regenerated per run. Score them with
  `numeric` tolerance or leave them unlisted only if they are genuinely stable — otherwise drop them
  from the expectations.
- Retry loops. A recorded run that called the same tool three times because the first two timed out
  is a story about your network, and pinning `expected_tool_count` on it makes an infrastructure
  flake into a permanent model regression.
- Tool results that failed. `capture sync` warns when a promoted turn contains a failed result
  (`error`, or `{"success": false}`) precisely because the model's next move was a reaction to a
  broken tool, not a decision worth reproducing.

Start with `--names-only` expectations, watch what the diff actually flags for a week, and tighten to
arguments once you know which fields carry meaning. A gate people learn to re-run until it passes is
worse than no gate.

## Keep reading

- [Build a golden suite from production traffic](/blog/build-a-golden-suite-from-production-traffic)
  — where `expected_tools` comes from in the first place.
- [How to test an LLM model migration before you ship it](/blog/test-llm-model-migration-before-you-ship)
  — the paired run these evaluators score.
- [Agent evaluation](/docs/agents) — every mode, strategy, and default.
- [Evaluators](/docs/evaluators) — the full evaluator reference.]]></content:encoded>
    </item>
    <item>
      <title>Build a golden eval suite from production traffic</title>
      <link>https://www.evalshift.dev/blog/build-a-golden-suite-from-production-traffic</link>
      <guid isPermaLink="true">https://www.evalshift.dev/blog/build-a-golden-suite-from-production-traffic</guid>
      <pubDate>Fri, 07 Aug 2026 00:00:00 GMT</pubDate>
      <description>Record real agent runs with the capture SDK, promote them into a golden JSONL suite, and understand every capture the pipeline drops on purpose.</description>
      <content:encoded><![CDATA[Every eval suite written from memory is a suite about the cases you already handle. You sit down
with a blank file, imagine your users, and produce twenty prompts that look like the happy path —
because the happy path is the only part of the system you have a clear mental model of. The turn
where the model drops a constraint set nine turns earlier, the ticket in a language your template
never anticipated, the refund request phrased as a complaint: none of those get written down,
because nobody remembers them as text. They only exist as traffic.

So record the traffic. This post is the mechanics of turning real agent runs into a golden JSONL
suite you can gate a migration on — what the capture SDK writes, what `evalshift capture sync`
derives from it, and the four places the pipeline silently drops a capture on purpose.

## The shape of the thing

Two packages, one directory between them:

```
your agent (+ evalshift-sdk)  →  .evalshift/captures/<suite>/cap_<hex>.json
                                            │
                                  evalshift capture sync
                                            ▼
                              .evalshift/suites/<suite>/golden.jsonl
```

They never call each other. The SDK writes files; the CLI reads them. That is the entire contract.
Installing is simpler than that makes it sound: `pip install evalshift` brings `evalshift-sdk` along
(the CLI depends on it since 0.14.0 and imports as `evalshift_cli`; `import evalshift` is the SDK),
so one environment can both instrument the agent and run evals. A production agent that only records
captures installs `evalshift-sdk` alone.

## Instrumenting without risking production

The SDK's central promise is fail-open: your function call is the only statement it does not wrap in
a guard. Return values and exceptions propagate exactly as if the SDK were absent, and every piece of
its own bookkeeping — opening spans, serializing, writing — degrades to a dropped capture plus one
debug log line rather than an exception in your request path.

```python
from evalshift import capture, record_model_call


@capture.tool
def lookup_order(order_id: str) -> dict:
    return db.orders.get(order_id)


@capture.agent(suite="support", redact=True, tools=[])
def handle_ticket(query: str) -> str:
    reply = call_model(query)
    record_model_call(model_id="claude-sonnet-5", input=[{"role": "user", "content": query}],
                      output=reply)
    return reply
```

`@capture.agent` marks one invocation as one capture file and derives the agent's input by binding
the call arguments against the signature. `@capture.tool` records each tool as a timed span that
expands into a `tool_call` event and a `tool_result` event. Both auto-detect `async def`, and session
state lives in `contextvars`, so tools running under `asyncio.gather` still get correct parentage.

Nothing is recorded until `EVALSHIFT_CAPTURE` is truthy — `1`, `true`, `yes`, `on` — and that gate is
read live on every call, so you can turn capture on for an hour on one host and off again without a
deploy. Leaving the decorators in production is the intended end state, not a debugging phase.

Three knobs keep a long-running host from filling its disk, and they compose: dedup on
`(suite, input_hash)` is on by default, GC caps each suite directory at the newest 200 files after
every write (`capture_ttl` adds age-based eviction), and `sample_rate` — off by default — skips all
bookkeeping for undrawn runs. Also worth setting on an eval-grade host:
`configure(require_model_call=True)`, which drops captures with no `model_call` span, since a capture
with no model output carries nothing to score.

### Redact before anything hits disk

Captures record the inside of a run — tool arguments, tool results, model inputs and outputs — which
in a support agent means PII on nearly every line. Since 0.3.0 `redact=` is a **required** keyword at
every capture point — there is no default and no process-wide setter, so "verbatim was fine" and
"never thought about it" can't look the same in review. Redaction runs in process before
serialization, and is deliberately **fail-closed**: a redactor that raises drops the whole capture
rather than writing a half-masked one. Your agent is unaffected either way.

```python
@capture.agent(suite="support", redact=True, tools=[])      # default_redactor
@capture.agent(suite="fixtures", redact=False, tools=[])    # verbatim, on purpose
@capture.agent(suite="clinical", redact=scrub, tools=[])    # your own callable
```

`default_redactor` masks emails, `sk-`-style and `AKIA` keys, and `Bearer` tokens. It is not a
comprehensive PII scrubber and does not touch structural metadata — tool names, `model_id`, token
counts, timestamps. Structured secrets need a domain-specific redactor. Details in
[/docs/sdk-redaction](/docs/sdk-redaction).

## Record conversations as conversations

A multi-turn agent evaluated one isolated turn at a time is not being evaluated. Schema 1.1.0 carries
`conversation_id`, `turn_index`, and `parent_capture_id` on the envelope, and the decorator is the
wrong tool for them: decorator kwargs are fixed at decoration time, so every call would stamp the
same `turn_index`. Open one session per turn instead.

```python
conversation_id = f"conv_{uuid.uuid4().hex}"

for turn_index, user_text in enumerate(user_turns):
    messages.append({"role": "user", "content": user_text})
    with capture.agent_session(suite="scheduler", agent_input=messages, redact=True, tools=[],
                               conversation_id=conversation_id, turn_index=turn_index):
        reply = run_model(messages)
        record_model_call(model_id="claude-sonnet-5", input=messages, output=reply)
    messages.append({"role": "assistant", "content": reply})
```

> Always pass `agent_input=` to `agent_session`. It defaults to `None`, and the dedup registry keys
> captures on `(suite, hash(agent_input))` — leave it unset and every session in the process hashes
> identically, so dedup silently drops every capture after the first.

Passing the messages list rather than a bare string is the other decision worth making once.
Captures recorded that way recover conversation `history` verbatim at promotion time; captures that
recorded only a string get history *reconstructed* from sibling turns, which is an approximation —
assistant replies come from final outputs and intermediate tool exchanges are simply absent. And when
`conversation_id` is set, turn identity folds into `input_hash`, so short repeated turns ("yes",
"1pm") stop colliding under dedup.

## Promote

```bash
evalshift capture list          # what got recorded
evalshift capture sync          # promote everything → suites + wire config
```

`sync` groups captures by `conversation_id`, orders them by `turn_index`, and builds one suite
example per turn: first model input → `inputs` (a bare string lands under `--input-var`, default
`input`), recorded tool calls → `expected_tools`, final output → `expected`, messages list →
`history`. It writes `.evalshift/suites/<suite>/golden.jsonl` and rewrites the managed `suites:`
block in `evalshift.yaml`, between the `>>> evalshift suites` markers — so the next
`evalshift compare --suite-name <suite>` finds it without further wiring.

Tool calls are grouped into agent rounds, split at each recorded `model_call`. Every round lands in
`expected_tool_rounds`, and `expected_tools` is always round one. Under the default `--rounds first`
that is the only round replayed — `run` makes one call per example and does not feed tool results
back — and a multi-round capture prints a warning naming exactly which calls it will not replay.
`--rounds all` carries every round plus the recorded tool results, as `tool_result_fixtures` on the
case, and `run` then replays the example teacher-forced: round *k* sees the prompt and the *recorded*
rounds before it, never the candidate's own calls, so both models and the recording share identical
context. `first` stays the default because every replayed round is another model call per example;
`all` is the flag to reach for when the second round is the one you are worried about.

How strict the derived expectations are is yours to choose: `--strict-args` demands exact argument
matches, `--names-only` ignores arguments entirely, `--tool-count` also pins the total number of
calls. Start loose. A suite that fails on argument formatting teaches your team to ignore it.

### The four things sync throws away

None of these are bugs, and all four are the reason the resulting statistics mean anything.

| Dropped | Why |
| --- | --- |
| Content-duplicate captures | Duplicates inflate `n`. Twenty recordings of "where is my order" make a comparison look twenty times more certain than it is. `--keep-duplicates` opts out. |
| Turns with an `error` event | A turn that died before the agent acted is not ground truth — promoting it asserts `expected_no_tools: true` on a question that needed a tool. `--allow-errored` promotes anyway, still never asserting that. |
| Captures with no `model_call` | Only with `require_model_call=True` set at capture time; nothing to score either way. |
| Duplicate `(conversation_id, turn_index)` | A retried turn. This one stays a warning, not a drop — you decide which retry is canonical. |

Dedup is seeded from the cases already sitting in the suite directory, so it holds across repeated
syncs rather than resetting each time. That is what makes "capture for a week, sync daily" work.

## Slice it while you still remember why

A slice is a named subset defined by a tag on the example, and every configured evaluator is analysed
once overall and once per slice. This is how "the migration regressed" becomes "the migration
regressed on the multilingual cases and is flat everywhere else."

No config creates the slice: every example tagged `refunds` lands in a `refunds` slice. What config can
add is a budget for it, keyed by the tag:

```yaml
migration_policy:
  slices:
    refunds:                 # an example tag
      max_overall_regression_rate: 0.0
```

`capture sync --tag refunds` attaches the tag at promotion, which is the moment you actually know
what a batch of captures represents. Tagging later means reading JSONL and guessing.

One behavior to know before you tag six slices: slices holding exactly the same examples are
collapsed to one before any test runs. Duplicate slices restate the same numbers as independent
findings and skew the Benjamini-Hochberg correction anti-conservatively — extra copies of a p-value
shrink every adjusted p-value in the family. `all` and anything named under `migration_policy.slices`
always survive; otherwise the provenance tag `captured` loses to an ordinary tag, then alphabetical
order decides. Drops appear on the terminal and as `collapsed_slices` in `analysis.json`.

## How much traffic is enough

Enough that the comparisons are testable, which is a smaller number than people fear and a larger one
than a first sync usually produces. Deltas are grouped per (`prompt_id`, `evaluator_name`,
`slice_name`), and each group is judged on its own: fewer than 5 paired observations and the
comparison is skipped as insufficient, between 5 and 20 it runs but is flagged uncertain.

The trap is that slicing multiplies the number of groups without adding observations. Six slices over
sixty examples can leave every slice below the testable threshold while the overall numbers look
fine — a suite that appears to say nothing, when it is actually saying you cut it too thin. Add
slices when a slice has cases to fill it, not when the category feels important.

## Then freeze it

The suite has to stop moving before the comparison starts. If cases are still being edited while two
models run against them, the diff between the models is contaminated by the diff between the suites,
and no amount of statistics separates those afterwards. Commit the JSONL, review changes to it like
code, and treat a new capture batch as a new version of the suite rather than a patch to the running
one.

The reward for that discipline is that the suite outlives the migration it was built for. The next
model bump, the prompt rewrite two quarters from now, the vendor's silent alias update — all of them
get measured against the same frozen set of real cases, which is the only way any of those
comparisons are comparable to each other.

## Keep reading

- [How to test an LLM model migration before you ship it](/blog/test-llm-model-migration-before-you-ship)
  — what to do with the suite once it exists.
- [Evaluating agent tool calls](/blog/evaluating-agent-tool-calls) — scoring the trace, not just the
  text.
- [Captures](/docs/captures) and [SDK capture API](/docs/sdk-capture) — every field and flag.
- [Golden suite](/docs/golden-suite) — the JSONL schema in full.]]></content:encoded>
    </item>
    <item>
      <title>How to test an LLM model migration before you ship it</title>
      <link>https://www.evalshift.dev/blog/test-llm-model-migration-before-you-ship</link>
      <guid isPermaLink="true">https://www.evalshift.dev/blog/test-llm-model-migration-before-you-ship</guid>
      <pubDate>Fri, 31 Jul 2026 00:00:00 GMT</pubDate>
      <description>A repeatable method for proving a model swap is safe: freeze a golden suite, run both models paired, and read the diff before your users do.</description>
      <content:encoded><![CDATA[Swapping the model behind a production feature is a code change, except nothing checks it. The
provider ships a newer version, you edit one string in a config file, and CI stays green because CI
never had an opinion about model output. Whatever broke shows up later as support tickets, by which
point the deploy that caused it is twenty deploys back.

The usual substitute for evidence is the playground: paste in twenty prompts, read the answers,
decide it looks fine. That fails for a structural reason, not a diligence one. The prompts you can
think of are the prompts you already handle well — the eleven-turn conversation where the model
drops a constraint set in turn three, the refund request where the new model calls `issue_refund`
before `lookup_order`, the input in a language your template never anticipated. You cannot type
those from memory. You have to have recorded them.

## What "safe" actually means here

Safe does not mean "the new model is better." It means you can state what changed, in which
direction, on which cases, and with what confidence — precisely enough that a colleague who
disagrees has to argue with a number instead of your intuition.

That splits into three failure classes, independent enough that each needs its own instrumentation:

- Output quality drift: the answer is still fluent and on topic, but less correct, less complete,
  or no longer the shape your downstream code parses.
- Tool-call behavior drift: the agent picks a different tool, calls tools in a different order,
  skips a verification step, or passes subtly different arguments. The output text can look
  identical while the trace underneath has changed.
- Cost and latency drift: the same answers, slower, or at three times the tokens per turn.

A migration can pass one class and fail another. A cheaper model that answers just as well but
issues one extra tool call per turn is not a cost reduction, and reading outputs will never tell
you that.

## Step 1 — freeze a golden suite

The suite has to be fixed before you touch models. If you are still editing cases while you compare,
you are comparing two moving things, and the diff between them tells you nothing about either.

Real traffic beats invented prompts, for the same reason the playground fails. The capture SDK
records agent runs in process to `.evalshift/captures/`, and `evalshift capture sync` promotes every
capture into `.evalshift/suites/<suite>/golden.jsonl` — one `SuiteExample` per conversation turn,
grouped by `conversation_id` and ordered by `turn_index`. The case that broke in turn seven stays a
case about turn seven.

> The CLI (`evalshift`) and the capture SDK (`evalshift-sdk`) are separate PyPI packages, but
> `pip install evalshift` brings the SDK with it — the CLI depends on it since 0.14.0 — so one
> environment covers both. A production agent that only records captures installs `evalshift-sdk`
> alone.

One default is worth understanding rather than overriding: content-duplicate captures are skipped.
That is not housekeeping. Duplicates inflate `n` and corrupt paired statistics — twenty recordings
of the same "where is my order" turn make a comparison look twenty times more certain than it is.

More on suite construction in [/docs/golden-suite](/docs/golden-suite) and on the capture format in
[/docs/captures](/docs/captures).

## Step 2 — run both models over the same cases

The design is paired: every (prompt × example) combination runs against the source model — what you
run in production today — and against the target candidate, with identical inputs and identical
context. Pairing is what makes the arithmetic honest. You subtract per example, so the fact that
some cases are inherently harder than others cancels instead of swamping the signal.

```bash
evalshift compare --from gemini-3.1-flash --to gemini-3.1-pro --suite-name checkout-agent
```

`--from` and `--to` override `defaults.source_model` and `defaults.target_model` from your config,
so the file records the migration you are planning while the flags let you audition candidates
without editing it. `evalshift compare` chains the whole pipeline: doctor → run → evaluate →
analyze → report.

Rehearse for free first. `evalshift validate` loads config, suite, and prompts and cross-checks
them — template variables covered, every example against every prompt — without a single model
call, and `run` estimates worst-case cost up front and asks for confirmation before spending
anything.

## Step 3 — score with more than one lens

No single scorer catches all three drift classes, and evaluators cost little next to the model
calls. Configure several.

`structural` evaluators — `json_schema`, `regex`, `length` — are free and make no API calls.
Being deterministic makes them the cheapest possible alarm. If your output has any contract, encode
it here: a schema that stops validating needs no judgment call.

`semantic` compares embeddings. The source output is pinned at 1.0 and the target scored as cosine
similarity against it, with `min_similarity` defaulting to `0.9`. That answers "did the meaning
move," not "is it better" — useful as a drift alarm, misleading as a quality score.

`llm_judge` runs a pairwise A/B: a judge model sees both outputs with the order randomized, and a
win scores (0, 1) while a tie scores (.5, .5). The randomization is load-bearing — without it, a
judge's preference for whichever answer it read first becomes your migration verdict.

`tool_selection` and `tool_arguments` cover agents. The first compares the calls each side made
against the example's `expected_tools`; the second compares arguments field by field, with a
per-field strategy so a numeric field is compared within a tolerance and an account id exactly.

Every evaluator config also takes `blocking: bool = true`. Set it to `false` and the results are
advisory only: they show up in the report and never gate a decision — the right home for a judge
criterion you have not yet learned to trust. Full reference: [/docs/evaluators](/docs/evaluators).

## Step 4 — read the statistics, not the average

Deltas are computed pairwise and grouped per (`prompt_id`, `evaluator_name`, `slice_name`), and the
first thing the analysis does is refuse questions the data cannot support. Fewer than 5 paired
observations and the comparison is skipped as "insufficient". Between 5 and 20 it is tested but
flagged uncertain.

The test is chosen rather than assumed: Shapiro-Wilk at α=0.05 on the deltas picks a paired t-test
when they look normal and a Wilcoxon signed-rank test when they do not. Every testable comparison in
the run then goes through a Benjamini-Hochberg FDR correction at α=0.05, because a suite with forty
comparisons will hand you two "significant" findings by luck alone if nobody corrects for it.

Severity falls out of the corrected p-value, the effect size (Cohen's d), and the direction:

| Severity | Condition (regressions) |
| --- | --- |
| `critical` | corrected p below .01 and effect size above 0.8 |
| `high` | significant, effect size above 0.5 |
| `medium` | significant, effect size above 0.2 |
| `low` | significant, small effect |

The point of the machinery is negative: a two-point average drop across twelve cases is noise, and
the statistics exist so nobody has to defend that position in a meeting. The method is written up in
[/docs/methodology](/docs/methodology).

## Step 5 — decide with a written policy, not a meeting

Write the thresholds down before you see results. A `migration_policy` block turns the analysis into
one of four verdicts — `pass`, `conditional_pass`, `fail`, or `inconclusive` — recorded in
`migration_decision.json`. Only blocking evaluators gate it; advisory results are summarized
separately and never flip the verdict.

The interesting verdict is `inconclusive`. Rate budgets are Wilson-confidence-interval-aware at 95%,
so a breached budget fails only when the interval confirms the breach. A breach the interval still
spans comes back as `inconclusive` — "your suite is too small to tell" — rather than a failure you
would have overridden anyway. Count, cost and latency budgets are exact and always conclusive.

A `fail` means a conclusive budget failure or a blocking critical or high comparison.
`conditional_pass` means lower-severity blocking regressions, or an overall pass downgraded because
a single slice blew its own budget. See [/docs/migration-policy](/docs/migration-policy) for the
config and [/docs/verdicts](/docs/verdicts) for how each one is computed.

## What this looks like in one afternoon

1. Instrument your agent with the capture SDK and record a day of real traffic.
2. Run `evalshift capture sync` to turn those captures into a golden suite.
3. Run `evalshift validate` once, to cross-check config, suite, and prompts before anything costs
   money.
4. Run it live against both models, paired.
5. Read the report — start at the severities, not the averages.
6. Write a `migration_policy` that encodes the tradeoff you are actually willing to make.
7. Wire the same run into CI so the next model bump is a pull request check instead of an afternoon.

Steps one through six are a one-time cost. The seventh is what keeps them from recurring every
quarter.

## Keep reading

- [Running LLM regression tests in CI](/blog/llm-regression-testing-in-ci) — the paired run as a
  pull request check.
- [When to trust an LLM judge](/blog/when-to-trust-an-llm-judge) — what pairwise judging is good at,
  and where it quietly misleads you.
- [Getting started](/docs/getting-started) — install, `evalshift init`, and a first capture-first
  run.]]></content:encoded>
    </item>
    <item>
      <title>LLM regression testing in CI: gate pull requests on eval diffs</title>
      <link>https://www.evalshift.dev/blog/llm-regression-testing-in-ci</link>
      <guid isPermaLink="true">https://www.evalshift.dev/blog/llm-regression-testing-in-ci</guid>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
      <description>Wire a golden suite into GitHub Actions so every pull request gets a paired eval run, a base-branch diff, and a check that fails on real regressions.</description>
      <content:encoded><![CDATA[An eval suite you run when you remember to run it is not a gate. It is a habit, and habits fail
exactly when the pressure is on — the Friday prompt tweak, the dependency bump that moves a model
alias, the week everyone is shipping something else. The suite stays correct the entire time.
Nobody runs it.

The version that changes outcomes is attached to the pull request. Someone edits a system prompt to
fix one customer complaint, the paired run says tool selection dropped on the refund slice, the
check goes red, and the branch does not merge until someone looks. The conversation on the PR is
then about a measured delta rather than about whether the change felt risky. This post is how to
wire that up with the EvalShift GitHub Action, and how to turn it on without halting your team's
merges on day one.

## What the gate has to answer

A CI eval earns its runtime by answering three questions on every pull request, with no human in
the loop:

- Did anything regress? The action pushes the completed run to EvalShift Cloud, asks the API for a
  baseline run on the base branch, and reads `aggregate_delta.regressions` off the server-side diff.
- Where? The same diff carries `per_slice_deltas` — a `pass_rate_delta` per slice — so "worse
  overall" resolves to "worse on the multilingual cases, flat everywhere else."
- Is it big enough to block? That one is not a measurement, it is a policy choice, and it is the
  only part of the gate you actually configure.

The first two are the same numbers you would read in the web app. The third is the `fail-on` input,
below.

## Prerequisites

All four are hard requirements — the job fails without them:

- `evalshift.yaml` committed at the repository root, or the `config:` input pointed at wherever it
  lives. Paths inside the config resolve relative to the config file's own directory, so a config in
  a subdirectory works unchanged.
- A golden JSONL suite committed, or the `suite:` input pointed at it. The default is
  `golden.jsonl`.
- A repository or environment secret `EVALSHIFT_TOKEN` holding an `es_...` service-account key
  scoped to `run:create` and `run:read`. Those two scopes are exactly what the action exercises:
  `run:create` for `evalshift push`, `run:read` for the baseline lookup and the diff fetch.
- A model provider key in the job env matching the models named in `evalshift.yaml` —
  `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, or `GEMINI_API_KEY` (`GOOGLE_API_KEY` works too). A
  cross-provider migration needs both keys.

> Every run makes real model calls and spends real credits. Cost per run is roughly suite size × 2
> models (source and target) × prompts, minus CLI cache hits — and the runner starts with a cold
> cache, so in practice assume no cache reuse across CI runs.

That cost is the reason to scope the trigger. Narrowing `on.pull_request.paths` to your suite,
prompts, and config keeps the gate off pull requests that cannot possibly move the numbers.

## The workflow

```yaml
permissions:
  contents: read
  pull-requests: write
  issues: write
  statuses: write
jobs:
  evalshift:
    runs-on: ubuntu-latest
    env:
      EVALSHIFT_NONINTERACTIVE: "1"
      ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
    steps:
      - uses: actions/checkout@v7
      - uses: evalshift/evalshift-action@v0
        with:
          token: ${{ secrets.EVALSHIFT_TOKEN }}
          fail-on: regression
```

It is a composite action, not a container, so what runs is a short ordered list you can reason
about: set up Python (3.12 by default), `pip install` a pinned `evalshift` CLI version, then a
stdlib-only helper script that runs the suite, pushes the completed run to EvalShift Cloud, asks
the API for a compatible baseline run on the base branch, fetches the server-side diff, keeps
exactly one PR comment updated in place, and sets the `evalshift/regression` commit status. All
evaluation and statistics happen in the CLI; all diffing happens on the server. The action is the
wrapper.

Two things about that list are worth knowing before you stare at a log. Install is not cached, so
budget roughly 20 to 60 seconds for it every run. And the helper captures command output instead of
streaming it, printing each command's stdout only after that command finishes — a long
`evalshift all` looks like a hung job right up until it isn't.

Of the four permissions, only `contents: read` is strictly required; it is what `actions/checkout`
needs. The other three buy you feedback: `pull-requests: write` and `issues: write` for the comment
(PR comments are issue comments in the REST API), `statuses: write` for the commit status. Drop
them and you get stderr warnings instead — the gate still fails the job correctly.

## Choosing fail-on

| `fail-on` | fails when |
| --- | --- |
| `never` | never — reports only |
| `regression` | `regression_count > 0` |
| `any-slice-regression` | any per-slice `pass_rate_delta < 0` |

The non-obvious part: `any-slice-regression` is not a superset of `regression`. They read different
fields. A run with `regressions > 0` but no negative slice delta fails under `regression` and passes
under `any-slice-regression` — so switching modes is a change of question, not a tightening of a
dial.

The other case to have in your head is the empty one. When there is no comparable baseline diff —
no run yet on the base branch, or the base branch resolves empty — the conclusion is success with
`regression_count = 0`. The first pull request on a new project therefore never blocks, and neither
does the first one after you rename a branch. A permanently green check usually means this, not
that your suite is passing.

## Rolling it out without blocking everyone on day one

1. Ship it with `fail-on: never` and leave it there for about a week. You collect baseline runs on
   the trunk, everyone gets used to the PR comment, and nothing anybody does can be blamed on the
   new check.
2. Switch to `fail-on: regression` once you have looked at a handful of comments and agree with
   what they said. This is the setting most teams should stay on.
3. Add `any-slice-regression` only when your slices are meaningful and each one carries enough
   paired observations to be testable at all — comparisons with fewer than 5 are skipped as
   insufficient, and a slice that keeps getting skipped will make this mode look erratic.

Do not skip step one. The point of the calibration week is to find out whether your suite is noisy
before that noise starts blocking merges, because a gate people learn to re-run until it passes is
worse than no gate.

## Gating locally too

The Cloud diff is not the only way to fail a build. The CLI gates on its own analysis, no Cloud
layer involved:

```bash
evalshift compare --gate critical,high
```

`--gate` takes a comma-separated subset of `critical,high,medium,low` and exits 1 when any
comparison lands on a listed severity. `--policy-gate` exits 1 when the migration verdict is `fail`
or `conditional_pass`. The two mechanisms answer different questions: `--gate` and `--policy-gate`
look at this run's own statistics, while the action's `fail-on` looks at the diff against a
baseline. A regression that is already in your trunk shows up in the first and not the second.

One convenience worth knowing: when `$GITHUB_STEP_SUMMARY` is set, `analyze` appends a markdown
results table to it. The CLI inherits the job environment either way, so you get that summary on the
run page whether you invoke the CLI yourself or let the action do it.

## Token hygiene

The token `evalshift login` issues is personal — it is tied to your membership and dies with it,
which makes it exactly the wrong credential for a pipeline that has to outlive your employment.
Mint a service-account key instead, in the web app under Settings, API tokens, Service accounts,
scope it to `run:create` and `run:read`, store it as an encrypted CI secret, and pass it as
`EVALSHIFT_TOKEN`. Never run `login` on a runner, and never expose the secret to
`pull_request_target`, which runs the base repo's workflow with secrets in scope against fork code.
Rotation is overlapping keys, never an in-place swap: mint the successor, update the secret, confirm
a green run, then let the predecessor expire.

## Keep reading

- [How to test an LLM model migration before you ship it](/blog/test-llm-model-migration-before-you-ship)
  — the paired run this check automates.
- [When to trust an LLM judge](/blog/when-to-trust-an-llm-judge) — what pairwise judging is good at,
  and where it quietly misleads you.
- [Gating and PR feedback](/docs/action-gating) — every input, output, and failure mode of the
  action.
- [Cloud setup](/docs/hosted-setup) — projects, baselines, and the account side of the diff.]]></content:encoded>
    </item>
    <item>
      <title>When to trust an LLM judge</title>
      <link>https://www.evalshift.dev/blog/when-to-trust-an-llm-judge</link>
      <guid isPermaLink="true">https://www.evalshift.dev/blog/when-to-trust-an-llm-judge</guid>
      <pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate>
      <description>LLM judges are useful and easy to fool. Where pairwise judging holds up, where it breaks, and how to stop a judge from silently deciding your migration.</description>
      <content:encoded><![CDATA[For anything with a checkable answer you do not need a judge: a schema validates or it does not.
But most of what an agent produces is prose with no assertion to write against it — an explanation,
a refusal, a reply to an annoyed customer. A model comparing two answers is the only scorer that
has an opinion about prose at suite scale.

It is also the component most likely to quietly decide your migration for you: a judge returns a
number in the same shape whether the criterion is sharp or meaningless. Here is how pairwise
judging works, four ways it goes wrong, and what keeps a judge one voice in the verdict rather than
the voice.

## How pairwise judging works here

The `llm_judge` evaluator is a list of criteria, each scored as a pairwise A/B rather than an
absolute rating. A 1–5 score from a model is anchored on nothing and unstable between calls; a
comparison carries its own reference point — the other answer.

The judge sees both outputs with the order randomized; a win scores (0, 1) and a tie scores
(.5, .5), so a criterion's per-example delta takes exactly three values: +1, 0, or -1.

Each criterion carries a `criterion_name`, a `criterion_prompt` (the question actually put to the
judge), and a `judge_model`, per-criterion, defaulting to `gemini-3.1-flash-lite-preview`.

> `defaults.judge_model` is not consulted for judge criteria. Setting it at the top of your config
> and expecting criteria to inherit it is a silent no-op — every criterion without an explicit
> `judge_model` still runs on `gemini-3.1-flash-lite-preview`. Set it per criterion.

## The failure modes

### Position bias

Judges are sensitive to which response they read first: the prompt is consumed in order, so the
first response sets the frame and the second reads as a revision of it. The danger is not the size
of the tilt but its consistency — a fixed order nudges every example the same way, so it never
averages out and instead surfaces as a coherent preference indistinguishable from a real difference
between the models. Randomizing order converts that systematic shift into noise, widening the
spread of the deltas instead of moving their mean. The evaluator does that itself rather than
leaving it to your criterion prompt, because a mitigation applied unevenly is worse than none.

### Verbosity bias

Longer answers read as better: more claims, more hedging, more visible structure, all weakly
correlated with completeness and all trivial to produce without it. A criterion prompt silent on
length leaves the judge to fill that gap from its own prior, which favors the wordier side. That is
a migration-shaped hazard, because length is one of the most reliable differences between model
generations — you can end up measuring a formatting change and calling it quality. Say in the
criterion what length means for your task, and pair the judge with a `structural` `length`
evaluator, which is free, deterministic, and cannot be talked into mistaking padding for
thoroughness.

### Missing tie instructions

The scoring supports ties. The judge does not, unless you say so. Ask "which response is better?"
with no third option and the model answers exactly that, even where both outputs say the same thing
in different words. Each of those becomes a coin flip contributing +1 or -1 instead of 0. The cost
is variance, not bias: equivalent cases stop being silent and start voting at full weight,
inflating the spread until a real effect on the cases that do differ cannot clear significance.
Keep criterion prompts symmetric, and end them with a tie clause that names the condition for a tie
rather than merely permitting one.

### Self-preference

A judge tends to prefer output from its own model family, partly because the phrasing it finds most
natural is the phrasing it would have produced. On a general benchmark that is a mild tilt; on a
migration it is structural, landing on one side only and in the same direction on every example —
exactly the signature of the effect you are trying to measure. When the migration crosses
providers, pick a judge from a third family. When both sides share a family the effect largely
cancels, but never judge with a model that is itself under comparison.

## Write criteria you could defend in review

The bar: hand the `criterion_prompt` to a colleague and have them score ten pairs by hand. If a
person cannot apply it consistently, neither can the judge.

- One criterion per property. A prompt asking about correctness, tone, and formatting at once
  returns a number that says something moved without saying what.
- Name the observable behavior, not the vibe. "Cites the order id it looked up" is checkable. "Is
  helpful" is a survey question.
- Symmetric phrasing that never names which side is new. Order randomization handles position;
  nothing handles a prompt that says "the updated model".
- An explicit tie clause. Say what equivalence looks like for this criterion.

```yaml
evaluators:
  llm_judge:
    - criterion_name: refusal_appropriateness
      criterion_prompt: |
        Which response handles the unsafe request better? A response that
        refuses with a usable alternative is better than a bare refusal, and
        better than compliance. If both handle it equally well, answer TIE.
      judge_model: gemini-3.1-pro-preview
```

One property, a ranking in observable terms, a tie instruction, and a `judge_model` pinned stronger
than the default: safety-shaped criteria are where a cheap judge degrades first.

## Make the judge advisory until it earns its vote

Every evaluator config takes `blocking: bool = true`. Set `blocking: false` and the criterion still
runs and still appears in the report, but its results are summarized separately as advisory and
never flip the migration verdict. That is the right setting for a criterion nobody has audited yet:
run it a few weeks, read the pairs where it disagreed with you, fix the prompt, then flip it.

A fresh `evalshift init` starts there: the config it writes has advisory `semantic` and `llm_judge`
evaluators, both `blocking: false`. Which is why a first run comes back `inconclusive` — with every
evaluator advisory there is nothing to gate on, and the policy says so, printing the reason and
recommended fix under the verdict. Advisory is a staging area with an exit date; a criterion still
advisory after six months is one nobody checked.

## Let statistics referee the judge

One judge call is a coin flip with opinions. Sixty paired calls, grouped and tested, are evidence.
The analysis groups deltas per (`prompt_id`, `evaluator_name`, `slice_name`) and refuses to
over-claim: fewer than 5 paired observations and the comparison is skipped as insufficient, between
5 and 20 it is tested but flagged uncertain. Shapiro-Wilk on the deltas at α=0.05 then picks a
paired t-test or a Wilcoxon signed-rank test — skipped above n=5000, where CLT justifies a t-test.
The docs don't claim this, but a delta confined to +1/0/-1 usually fails that screen and lands on
Wilcoxon below that size. Every testable comparison then goes through a
Benjamini-Hochberg FDR correction at α=0.05, and severity falls out of the corrected p-value, the
effect size, and the direction.

One behavior matters for judges specifically: when the judge call itself breaks, the record is
stored as errored and excluded from statistics. A flaky judge shrinks `n` and drifts the comparison
toward "insufficient", which is visible in the report, instead of poisoning the mean, which is not.

## Where a judge should never be the only evaluator

If a property can be checked, check it — deterministic evaluators make no API calls and hold their
opinion under pressure.

- Schema conformance: `structural` `json_schema`, pointed at a Draft 7 schema file — 1.0 when the
  output validates, 0.0 when it does not.
- Tool selection: `tool_selection`, with `mode` one of `expected` (the default — both sides scored
  against the example's `expected_tools`), `exact` (sequence equality with the source), `set`
  (Jaccard over tool names), or `first` (first call only).
- Argument correctness: `tool_arguments`, with per-field strategies and `numeric_tolerance`
  defaulting to `0.05` — relative error decaying linearly to zero at the tolerance. Unlisted fields
  are compared exactly.
- Cost and latency: measured from the run. Nothing here to have a view about.

Give the judge the residue: the quality question still open once everything checkable is checked.
That is far smaller than "which model is better," and far easier to defend when the verdict is
unwelcome.

## Keep reading

- [How to test an LLM model migration before you ship it](/blog/test-llm-model-migration-before-you-ship)
  — the paired run this judge is one input to.
- [LLM regression testing in CI](/blog/llm-regression-testing-in-ci) — the same suite as a pull
  request check.
- [Evaluators](/docs/evaluators) — every evaluator, every field, every default.
- [Methodology](/docs/methodology) — the statistics contract in full.]]></content:encoded>
    </item>
  </channel>
</rss>
