LLM migration testing for AI agents. Don't switch models on vibes.

Read the diff. Ship with proof.

Switch models or providers without guessing whether your agent will break. EvalShift replays captured agent traces on your current model and the candidate, scores outputs, tool calls, cost and latency, and tells you whether the change is a real regression.

● local-first

Your eval data stays on your machine unless you explicitly push a run to EvalShift Cloud.

Works with Anthropic, OpenAI, Google and DeepSeek models.

Migration verdictmain_chat · gemini-3.7-flash → gemini-3.5-flash-lite
16 captured examples replayed on both models
FAIL5 of 7 budgets within policy
Confirms actions it never performedAsked to book coffee, it says it's on the calendar — without calling any tool.
Routes to the wrong toolA weather question gains a display_info call; the answer becomes “it's on your screen”.
Searches instead of answeringAsked for an opinion, it issues search_web and returns no text at all.
RecommendationDo not migrate. −58.8% cost does not cover 25% tool divergence.
Cost−58.8%
Latency−75.2%
Tool divergence25.0%ceiling 10%
Equivalence61.4%floor 75%

How it works

Compare your current and candidate models on real production behavior. Catch regressions before they reach users.

captured turn · main_chat

“Cool. What's your take on whether working late is worth it?”

recorded by @capture.agent · turn 1
Source tracescore 1.000
gemini-3.7-flash(no tool calls)

final text: It's occasionally necessary during crunch times or when you're in a great flow state, but making it a habit usually costs more in focus and recovery than it delivers in output. Sustainable momentum during your core hours almost always wins out long term.

Target tracescore 0.000
gemini-3.5-flash-lite

1. search_web ({"query": "is working late worth it productivity well-being"})

final text: (none — the model returned no text)

Why flaggedRouting — divergenceTOOL_SELECTION_DRIFT

The source called no tools; the target called search_web.

Tool divergence 25.0% against a 10% ceiling: FAIL.

One decorator. The SDK records what your agent really does — prompts, tool calls, arguments — while it runs.

Replay those cases on the model you have and the model you're switching to. Same inputs, both models, scored side by side.

A verdict — pass, conditional, or fail — with the reasons, the numbers and the cost difference.

It tells you what broke, in plain English.

The same run as the report card above. Every figure is measured; the prose is generated from the figures, never the other way round.

Score change by evaluator

16 examples · target minus source

3 of 5 evaluators regressed

Each evaluator measures quality a different way. The findings describe what the regressions look like in the examples.

What changed
Confirms actions it never performedAsked to book coffee with Marcus, the target replies that it is on the calendar without calling any tool. Asked to start a groceries list, it creates the list, skips add_to_list, and still reports the three items as added.
Routes to the wrong toolA weather question gains a display_info call and the answer becomes “it's on your screen” instead of the spoken forecast. A café request drops the opening-hours search and adds an unrequested add_note.
Searches instead of answeringAsked for an opinion on working late, the target model issues a search_web call and returns no text at all, where the source model answered in two sentences.
RecommendationDo not migrate to gemini/gemini-3.5-flash-lite. A −58.8% cost saving does not cover a tool divergence of 25.0%: fix the 5 tool-selection drift and 7 semantic regression cases before re-evaluating.

Runs on your laptop. Blocks the merge in CI.

The Action runs your golden suite on every pull request, keeps one comment updated, and fails the check when your migration policy says the candidate is not safe to ship. Nobody has to remember to look.

On your laptop
~/repos/butler_service
evalshift compare


Migration verdict: fail

  Do not migrate globally under the configured policy.


run    r_20260827_main_chat_2b6679
suite  main_chat
report .evalshift/runs/r_20260827_main_chat_2b6679/report.html
On the pull request
pull request · swap gemini-3.7-flash for gemini-3.5-flash-lite● checks failed
  • buildSuccessful
  • evalshift/regression2 budgets breached

Merging is blocked

The evalshift/regression check failed.

Four lines in your workflow
- uses: evalshift/evalshift-action@v0
  with:
    token: ${{ secrets.EVALSHIFT_TOKEN }}
    fail-on: policy

When it's safe, it says so.

  • evalshift/regression10 / 10 budgets within policy

A pass is not silence. The report still lists what moved, so a non-blocking drop in similarity is a note to read before promoting, not a surprise later.

How gating decides →

Your first comparison, in three steps.

The CLI, the capture SDK and the HTML report are open source. No account, and nothing leaves your machine.

  1. Install the CLI

    One package installs the CLI and the capture SDK. evalshift init adds a minimal evalshift.yaml to your repo.

    uv pip install evalshift
    evalshift init
  2. Capture your agent

    Decorate the entry point and run it as usual. Nothing is recorded unless EVALSHIFT_CAPTURE is set, so the decorator can stay in production.

    from evalshift import capture
    
    @capture.agent(suite="support_agent", redact=True, tools=[])
    def handle(message: str) -> str: ...
    
    EVALSHIFT_CAPTURE=1 python agent.py
  3. Compare against the candidate

    Name the model you're switching to as target_model in evalshift.yaml, then turn the captures into a golden suite and replay it.

    evalshift capture sync
    evalshift compare --suite-name support_agent
Getting started guide →

Want shared run history and pull-request gating? Start free on EvalShift Cloud, no card needed, or compare plans on the pricing page.

Python 3.11+ · Apache-2.0

Release notes, by email.

A short email when the CLI, SDK or GitHub Action ships something new. Unsubscribe anytime.