LLM migration testing for AI agents. Don't switch models on vibes.
Read the diff. Ship with proof.
Switch models or providers without guessing whether your agent will break. EvalShift replays captured agent traces on your current model and the candidate, scores outputs, tool calls, cost and latency, and tells you whether the change is a real regression.
Your eval data stays on your machine unless you explicitly push a run to EvalShift Cloud.
Works with Anthropic, OpenAI, Google and DeepSeek models.
How it works
Compare your current and candidate models on real production behavior. Catch regressions before they reach users.
“Cool. What's your take on whether working late is worth it?”
recorded by @capture.agent · turn 1final text: It's occasionally necessary during crunch times or when you're in a great flow state, but making it a habit usually costs more in focus and recovery than it delivers in output. Sustainable momentum during your core hours almost always wins out long term.
1. search_web ({"query": "is working late worth it productivity well-being"})
final text: (none — the model returned no text)
The source called no tools; the target called search_web.
Tool divergence 25.0% against a 10% ceiling: FAIL.
One decorator. The SDK records what your agent really does — prompts, tool calls, arguments — while it runs.
Replay those cases on the model you have and the model you're switching to. Same inputs, both models, scored side by side.
A verdict — pass, conditional, or fail — with the reasons, the numbers and the cost difference.
It tells you what broke, in plain English.
The same run as the report card above. Every figure is measured; the prose is generated from the figures, never the other way round.
Score change by evaluator
16 examples · target minus source3 of 5 evaluators regressed
Each evaluator measures quality a different way. The findings describe what the regressions look like in the examples.
Runs on your laptop. Blocks the merge in CI.
The Action runs your golden suite on every pull request, keeps one comment updated, and fails the check when your migration policy says the candidate is not safe to ship. Nobody has to remember to look.
evalshift compare Migration verdict: fail Do not migrate globally under the configured policy. run r_20260827_main_chat_2b6679 suite main_chat report .evalshift/runs/ r_20260827_main_chat_2b6679/ report.html
- buildSuccessful
- evalshift/regression2 budgets breached
Merging is blocked
The evalshift/regression check failed.
- uses: evalshift/evalshift-action@v0
with:
token: ${{ secrets.EVALSHIFT_TOKEN }}
fail-on: policyWhen it's safe, it says so.
- evalshift/regression10 / 10 budgets within policy
A pass is not silence. The report still lists what moved, so a non-blocking drop in similarity is a note to read before promoting, not a surprise later.
How gating decides →Your first comparison, in three steps.
The CLI, the capture SDK and the HTML report are open source. No account, and nothing leaves your machine.
Install the CLI
One package installs the CLI and the capture SDK.
evalshift initadds a minimal evalshift.yaml to your repo.uv pip install evalshift evalshift init
Capture your agent
Decorate the entry point and run it as usual. Nothing is recorded unless
EVALSHIFT_CAPTUREis set, so the decorator can stay in production.from evalshift import capture @capture.agent(suite="support_agent", redact=True, tools=[]) def handle(message: str) -> str: ... EVALSHIFT_CAPTURE=1 python agent.py
Compare against the candidate
Name the model you're switching to as
target_modelin evalshift.yaml, then turn the captures into a golden suite and replay it.evalshift capture sync evalshift compare --suite-name support_agent
Want shared run history and pull-request gating? Start free on EvalShift Cloud, no card needed, or compare plans on the pricing page.
Python 3.11+ · Apache-2.0
