The EvalShift GitHub Action turns your golden suite into a merge gate. On every pull request it runs the suite against both models, pushes the result to EvalShift Cloud, compares it to the latest run on your base branch, and — by default — fails the check when your migration policy says the candidate is not safe to ship. It is a composite action that installs Python and the pinned CLI — the evalshift-version default in action.yml, 1.1.0 at the time of writing, is bumped by the action’s pin workflow after each CLI release — nothing compiled, nothing containerised. What it adds on top of the CLI: a plan preflight, Cloud push, baseline lookup, cross-branch diff, the policy verdict, one self-updating PR comment, a commit status, and an exit code.
The Action is open source under MIT — the code lives at evalshift/evalshift-action ↗.
evalshift compare --gate gives you local-only CI gating for free. But the full setup we recommend for teams is all three — SDK capture, CLI evaluation, and the Action as the PR gate. See the recommended workflow.## Who this is for
You already have a golden suite and you run it by hand before shipping a model or prompt change. That works right up until it doesn’t: someone edits a system prompt on a Friday, nobody re-runs the suite, and the regression ships. The Action closes that gap — it makes “did this change make the model worse?” a required check, answered by the same statistics you’d get locally, on a pull request, before anyone can merge.
If you don’t have a suite yet, start with the CLI — Getting started gives you a working project in one command. Come back here once evalshift compare passes locally.
## Prerequisites
| Requirement | How to get it |
|---|---|
| evalshift.yaml committed | evalshift init writes a capture-first starter |
| A golden JSONL suite committed | evalshift capture sync writes one from recorded production captures (or hand-write golden.jsonl) |
| Repository secret EVALSHIFT_TOKEN | A service-account key, not a personal token: EvalShift Cloud → Settings → API tokens → Service accounts, role member, scopes run:create + run:read + policy:read. Starts with es_. See Cloud setup. |
| A migration_policy block in evalshift.yaml | What the default fail-on: policy gates on. evalshift init scaffolds one; without it the run is reported as ungated. See Migration policy. |
| A provider API key as a repository secret | Whichever provider your config's models belong to |
| Any repository, public or private | Private repositories are included on every plan, the free one included. The Action still reports the repository’s visibility on every run. See pricing. |
Verify locally first. If evalshift compare --yes doesn’t pass on your machine, it will not pass on a runner — you’ll just pay for the model calls to find out.
## Quick start
Create .github/workflows/evalshift.yml:
name: evalshift
on:
pull_request:
push:
branches: [main]
permissions:
contents: read
pull-requests: write
issues: write
statuses: write
jobs:
evalshift:
runs-on: ubuntu-latest
env:
EVALSHIFT_NONINTERACTIVE: "1"
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
steps:
- uses: actions/checkout@v7
- uses: evalshift/evalshift-action@v0
with:
token: ${{ secrets.EVALSHIFT_TOKEN }}
fail-on: policy # the default; gates on your migration policyevalshift init --ci scaffolds a fuller, per-suite version of this: a discover job that lists the suites under .evalshift/suites/, one job per suite selecting it with suite-name, a pinned evalshift-version, and an evalshift gate job that collapses them into the one check to require in branch protection.
Keep the push: branches: [main] trigger. The default fail-on: policy judges each run against its migration policy and needs no baseline, but the diff in the PR comment — and the diff-based regression and any-slice-regression modes — compare against the most recent run on your base branch. Without trunk runs, every PR reports “no baseline”, and the diff-based modes pass unconditionally.
push trigger record a baseline on main, and the next PR gets a real comparison.## Secrets and provider keys
The Action does not manage provider credentials. It passes the job environment through to the CLI unchanged, so set the key as a job-level env: entry (ANTHROPIC_API_KEY, OPENAI_API_KEY, GEMINI_API_KEY/GOOGLE_API_KEY, DEEPSEEK_API_KEY) and the CLI picks it up. Which key you need follows from defaults.source_model and defaults.target_model in your config — comparing across two providers means both keys:
env:
EVALSHIFT_NONINTERACTIVE: "1"
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}EVALSHIFT_NONINTERACTIVE: "1" is recommended. The Action already passes --yes, which skips the CLI’s cost confirmation, but the env var covers any other prompt — and a prompt on a runner means a hung job.
## Where to next
- +Inputs & outputs — every input, output, and permission, plus version pinning.
- +Gating & PR feedback — the
fail-onmodes, the policy verdict, the PR comment, and how baselines resolve. - +Cost control & recipes — run it only when it matters.
- +Cloud setup — the account and token the Action needs.
