// action · reference
Inputs & outputs
the full surface
Everything evalshift/evalshift-action@v0 accepts and produces. Boolean inputs accept 1, true, yes, on (case-insensitive); anything else is false.
## Inputs
| Input | Default | What it does |
|---|---|---|
| token | — (required) | EvalShift Cloud API token, an es_... value. Masked in logs and redacted from CLI output. |
| host | https://api.evalshift.dev | Cloud API base URL. Set only for a self-hosted or staging deployment. |
| config | evalshift.yaml | Path to your config, relative to the repository root. Paths inside the config resolve relative to the config file's own directory, so a config in a subdirectory works. |
| suite | golden.jsonl | Path to the golden JSONL suite, relative to the repository root. Selects a file only: a suite that wires its own evaluators: block under suites: must be selected with suite-name instead, or it is scored with the top-level evaluators. Mutually exclusive with suite-name; golden.jsonl is used when neither is set. |
| suite-name | — | Name of a suite wired under suites: in evalshift.yaml (what evalshift capture sync writes). Preferred over suite: it carries that suite's own evaluator block. Needs evalshift-version >= 0.14.0. |
| evalshift-version | 1.1.0 | Exact CLI version installed from PyPI. Pin this for run-to-run reproducibility across CLI releases. |
| python-version | 3.12 | Python used to install and run the CLI. Must satisfy the CLI's minimum (3.11 for 1.1.0). |
| fail-on | policy | Gating mode: policy, never, regression, or any-slice-regression. See Gating & PR feedback. |
| require-policy | false | Whether a run pushed without a migration_policy fails the job. By default such a run merges — it is reported as ungated, with a workflow warning and a commit status that says the gate is off. Only fail-on: policy consults this. |
| branch | auto | Candidate branch name recorded on the Cloud run. Auto-detected from the PR head ref, else the pushed ref. |
| base-branch | auto | Branch to look for a baseline run on. Auto-detected from the PR base ref, else the current ref. Resolving to empty means no baseline is fetched: the diff-based modes always pass, while fail-on: policy still gates on the verdict. |
| create-project | true | Whether evalshift push may auto-create the Cloud project when it doesn't exist. Set false to make a missing project a hard failure. |
| comment | true | Whether to create or update the PR comment. Set false to keep the commit status but stay out of the conversation. |
| github-token | github.token | Token used for the PR comment and the commit status. Override only to have a bot account post instead of github-actions. |
| repo-private | repository visibility | Whether this repository is private, used for the plan preflight. Defaults to github.event.repository.private; reported to EvalShift, not verified by it. Set it explicitly only if you mirror a private repository into a public one, or vice versa. |
## Outputs
| Output | Value |
|---|---|
| run_url | Cloud run URL for this run. |
| diff_url | Cloud diff URL comparing this run to the baseline. Empty string when no compatible baseline was found. |
| run_id | Hosted EvalShift run id — the server-minted id every /runs/{id} API route takes. Not the local r_... run directory name. |
| regression_count | Number of regressed examples in the Cloud diff. 0 when there is no baseline. |
| conclusion | success or failure, reflecting fail-on. A policy that declined to decide is a success here — read the PR comment or the commit status for the verdict itself. |
Consume them from a later step:
yaml
- uses: evalshift/evalshift-action@v0
id: evalshift
with:
token: ${{ secrets.EVALSHIFT_TOKEN }}
- run: echo "Cloud diff ${{ steps.evalshift.outputs.diff_url }}"Outputs are written before the comment and status calls, so they are still available even if the job lacks permission to comment.
## Permissions
| Permission | Why |
|---|---|
| contents: read | Checking out the repository. |
| pull-requests: write | Posting the PR comment. |
| issues: write | PR comments are issue comments in the GitHub API. |
| statuses: write | Setting the evalshift/regression commit status. |
Only contents: read is strictly required. If the comment or status permissions are missing, the Action logs a warning and carries on rather than failing the run — the gate still works.
## Versioning and stability
Pin to @v0 to track the latest v0.x, or to an exact tag such as @v0.5.1 for a fully reproducible workflow. The evalshift-version input pins the CLI separately — pin both if you want a workflow that behaves identically six months from now.
+
wrapper, not reimplementation
All evaluation and statistics happen in the CLI; all cross-branch diffing happens server-side, and under
fail-on: policy the verdict is the one the CLI computed against the pushed migration policy — the Action never re-implements a threshold. To understand what regression_count actually means, read the statistical methodology — it is paired statistics with Benjamini–Hochberg correction, not a threshold on an average.