By default the Action gates on the verdict EvalShift Cloud returns for the run under its migration policy; the other modes gate on the Cloud diff between this run and the latest compatible run on the base branch. Either way it reports through one PR comment and one commit status. This page covers the fail-on modes, what lands on the pull request, and how branches and baselines resolve.
## The fail-on modes
| Mode | The job fails when |
|---|---|
| policy | Default. Fails only when the policy verdict for this run is fail — the verdict the CLI computed against the migration_policy block of your evalshift.yaml, pushed with the run and returned by EvalShift Cloud. |
| never | Never fails. Records the run, pushes it, comments — but never blocks the merge. Use while you're still calibrating a suite. |
| regression | Fails when the Cloud diff reports one or more regressed examples in aggregate. |
| any-slice-regression | Fails when any slice's pass rate moved down, even when the aggregate is flat or improved. Catches one slice degrading while the overall number hides it. |
### policy — the governed gate
policy is the only mode that enforces what you configured. The CLI resolves the migration_policy block, evalshift push carries it with the run, and the Action calls GET /runs/{id}/policy-check and follows the answer — the same verdict your local evalshift compare and the run’s Policy tab show. That makes it disagree with regression in both directions, on purpose: a run with regressed examples that stays inside every budget passes, and a run with zero regressions that busts a cost, latency or per-slice budget fails. The policy answers with one of four statuses, and only one of them fails the job:
| Status | Gate | What it means |
|---|---|---|
| pass | merges | Every budget is within policy. |
| conditional_pass | merges | Every budget held and nothing critical or high regressed, but medium/low regressions or comparisons that scored zero pairs came with it. A pass with caveats — the comment says so and prints the server's reason. |
| fail | blocks | A budget was busted, or a critical/high regression is blocking. |
| inconclusive | merges | The policy could not decide — never rendered as a pass. Only the server's reason string tells the causes apart, so the Action prints it verbatim. |
| inconclusive, policy_source: none | merges (blocks with require-policy: true) | Nothing gates this PR: the run carried no migration policy. The Action prints a ::warning:: annotation, says the gate is off in the commit status, and asks for the fix in the PR comment. |
A status outside those four — something a newer server grew — is handled like inconclusive: it does not fail the job and is never rendered as a pass. require-policy governs only the no-policy row. If the policy check itself is unreachable, answers 404, or holds no decision for the run, the Action does not go quietly green: it falls back to regression gating for that run and says so in the log, the commit status and the PR comment.
### The diff-only modes
any-slice-regression is not simply “stricter than regression” in every case; it’s a different question. A run where the aggregate regression count is above zero but no individual slice moved down will fail under regression and pass under any-slice-regression. If you want both guarantees, run the Action twice with different modes (and comment: "false" on one of them), or keep regression and rely on slices for diagnosis rather than gating.
never for a week or two while the suite settles, then move to policy and tune the budgets in evalshift.yaml until the gate agrees with your own judgement of which PRs should have been blocked. Reach for regression or any-slice-regression only when you deliberately want a diff-shaped question the policy doesn’t ask.When no compatible baseline run exists on the base branch, there’s nothing to compare against: the diff-based modes pass, regression_count is 0, and the PR comment says so explicitly. policy still asks the server for a verdict — the policy is about this run, not about the comparison.
## What lands on the pull request
One comment, updated in place. The Action maintains exactly one comment per PR, marked with a hidden HTML marker so it edits itself on every push instead of stacking up. With a baseline present it shows the conclusion, the Cloud run link, the regression count, the diff link, the pass-rate movement, and up to five regressed slices, worst first. When nothing regressed you get a single No regressed slices row. Percentages are rounded for display only — gating uses the raw values, so a slice can appear as 0 pts and still count as a regression. Without a baseline, the slice table is replaced by a single line saying no compatible baseline run was found.
Under fail-on: policy the comment also carries the policy decision and its reason, a Policy budgets table (each budget’s scope, observed value, allowance and result, failing first, capped at 12 rows) and a Blocking regressions table (capped at 10). A blockquote states the caveat on a conditional_pass, and says plainly when the policy could not decide or the policy check could not be reached. The policy sections render with or without a baseline; the other modes never ask for a verdict, so they show none.
A commit status. Context evalshift/regression — the name is fixed in every mode, including policy — linking to the Cloud diff (or the run, when there’s no diff), with the gate’s one-line summary as its description. It is set on push events too, not just pull requests. With a single suite, this is what you add to branch protection; with several suites every job writes the same context, so require a join job instead — see Make it a required check.
## How it works, step by step
- +Install —
actions/setup-python, thenpip install evalshift==<version>. No pip caching, so budget roughly 20–60 seconds. - +Preflight — reads
project: <org>/<project>from your config and makes one call to EvalShift Cloud (POST /runs/preflight, authorized byrun:create) asking whether the org’s plan covers this job: subscription status, seat overage, the monthly run quota (runs_per_month), private-repo CI, and the declared matrix width — all checked before a single model call. A402stops the job here — even atfail-on: never— with an::error title=EvalShift::annotation, a## EvalShift did not runstep summary and, on a PR, the comment. Three more answers stop the job the same way — the annotation and the step summary, but no PR comment, and the job exits1: a401(the token itself is bad, expired, or revoked — not a permission problem), a403(the key lacksrun:create, named in the message), and a404withcreate-project: false(no project at that slug this key can reach). A404withcreate-project: trueis not a stop: it prints an::notice title=EvalShift preflight::— the first push creates the project — and the job continues. Against an older server that predates this endpoint (405), it warns (plan preflight skipped: hosted EvalShift at <host> predates POST /runs/preflight (HTTP 405)) and continues, with no fallback to the old two-call flow. Any other failure (an outage, a timeout, a malformed response) also warns (plan preflight skipped: …) and continues — the server still enforces plan limits on upload. Skipped entirely when the config has no top-levelproject:. - +Run —
evalshift all --yes --config …plus the suite selection (--suite-name …or--suite …), the full local pipeline. Artifacts land in.evalshift/runs/<run-id>/, includingreport.html.allis the permanent alias forcompare: the Action invokes it on purpose, because it is the one spelling every pinned CLI version understands. - +Push —
evalshift pushuploads the run bundle, creating the project ifcreate-projectallows it. Git metadata travels with the bundle so the server can pair this run with base-branch runs later. - +Find a baseline — asks the Cloud API for the latest compatible run on the base branch. “Compatible” is a server-side judgement — a suite that changed shape can’t be diffed against an older one.
- +Fetch the diff — aggregate and per-slice deltas from the Cloud API.
- +Ask the policy gate — under
fail-on: policy,GET /runs/{id}/policy-checkreturns the verdict on this run against the migration policy it was pushed with, plus the budget arithmetic behind it. Skipped in the other modes. - +Report — writes the five outputs, upserts the PR comment, sets the commit status.
- +Gate — exits non-zero when the gate says so: the policy verdict under
fail-on: policy, the diff otherwise.
## Branch and baseline resolution
The Action figures out two branch names, and both matter:
- +
branch— the candidate, recorded on the Cloud run. From the PR head ref, else the pushed ref. - +
base-branch— where to look for a baseline. From the PR base ref, else the current ref.
On a push event, base-branch falls back to the branch being pushed. A push to main therefore diffs against the previous main run. That’s intentional: it tracks trunk drift over time, and it’s how baselines get recorded in the first place. Override branch/base-branch only when your branch naming genuinely differs from your git refs. If base-branch resolves to an empty string, the Action skips the baseline lookup entirely: the diff-based modes always pass, while policy still gates on the verdict. Baselines and diffs themselves are covered in Baselines & diffs.
