## Security model
- +Secrets are masked and redacted. The Cloud token and GitHub token are registered with GitHub’s log masking before anything else runs. On top of that, the Action redacts any environment value whose name contains
TOKENorSECRET, or ends inAPI_KEY, out of the CLI’s stdout and stderr before printing it — so a CLI that echoes a key in an error message doesn’t leak it into your logs. - +Tokens never appear in argv. The Cloud token and host URL reach the CLI through the environment only, so they can’t show up in a process listing or a
command failed:message. - +Provider keys never leave the job. The Action passes them to the CLI and nowhere else. They are not uploaded to EvalShift Cloud, not written to the bundle, not sent to GitHub.
- +The Action never writes to your repository. No commits, no pushes, no file mutations outside
.evalshift/in the workspace. - +Dependencies. The runtime helper is stdlib-only, and
pip-auditruns in the Action repo’s own CI.
pull_request from a fork, so the Action will fail on token being empty. That’s GitHub’s design, and working around it with pull_request_target means running untrusted code with your secrets in scope — don’t, unless you fully understand the exposure.## Limits and known edges
- +The PR comment lookup reads only the first page of comments. On a very long PR thread the EvalShift comment can fall off page one, and a second comment gets created instead of the first being updated.
- +The comment marker and status context are global constants. Parallel invocations on the same PR overwrite each other. One commenting invocation per PR.
- +The local run id is the newest directory under
.evalshift/runs. If a step between the run and the push touches an older run directory’s mtime, the wrong run gets pushed. In a normal workflow this never happens. - +The Cloud run URL is parsed from CLI stdout, and the run id is read out of its
/app/{org}/{project}/runs/<uuid>path. A CLI release that changes how the push result is printed would break this; the Action repo’scli-contractCI job guards flag renames but not output shape. A URL the Action cannot read an id out of fails the step rather than gating on a guess. - +No retries on Cloud API calls. A 30-second timeout, one attempt. A transient Cloud outage fails the step rather than silently passing — deliberate, but it means a flaky network reads as a failed job.
- +
fail-ondecides the exit code, not whether the run happened. Even atnever, the run executes, costs money, and pushes. - +A policy that cannot decide does not fail the job.
inconclusive— and any status the Action doesn’t know — is reported loudly and gated as a non-failure. The single exception is a run that carried no policy at all, whichrequire-policy: trueturns into a failure. - +
conditional_passmerges. It is a pass by design. If medium/low regressions should block, tighten the policy — the Action will not second-guess the verdict. - +An unreachable policy check degrades to
regression. The Action says so in the log, the commit status and the PR comment, but for that run the gate is the diff, not your policy. - +A denied preflight fails the job at
fail-on: nevertoo.fail-ongoverns the verdict; a plan that doesn’t cover the run is a different question, and the run never happens.
## Troubleshooting
The log lines below quote evalshift all because that is what the Action invokes. all is the CLI’s permanent alias for compare and the one spelling every pinned CLI version answers to, including releases that predate the rename — so the Action stays on it deliberately. What you type locally is compare.
| Symptom | Cause and fix |
|---|---|
| input 'token' is required | The token: input is empty. Either the secret isn't set, or this is a fork PR where secrets aren't exposed. |
| command failed (1): evalshift all --yes ... | The CLI itself failed — bad config, missing provider key, model API error. The CLI's own (redacted) stderr is printed directly above this line. Reproduce with the same command locally. |
| no local EvalShift runs found in .../.evalshift/runs | evalshift all exited successfully but wrote nothing where the Action looks. Usually a config that redirects run artifacts elsewhere, or a working-directory mismatch. |
| evalshift push did not print a hosted run URL | The push didn't emit a URL on its last output line. Run evalshift push <run-id> locally against the same host and see what it prints. Also check for a CLI version mismatch. |
| could not read a server run id out of the hosted run URL: ... | The push printed a URL, but without an /app/<org>/<project>/runs/<uuid> path. Every Cloud call afterwards is keyed on that id, so the Action stops rather than gate on a guess. Usually an older CLI; check the printed URL in the step log and pin a newer evalshift-version. |
| HTTP 401 from the Cloud API | Bad, revoked, or expired EVALSHIFT_TOKEN, or the wrong host. Verify with evalshift whoami locally using the same token. A key past its rotation grace window authenticates as nobody. |
| The EvalShift token is missing the '<key>' permission | An HTTP 403: the key authenticated, but its scopes — or its service account's role — don't cover what the step needed. On the Action's calls, 403 means a missing permission, never a wrong project. Widen the scope or mint a key that holds it; the Action always needs run:create, run:read, and policy:read. |
| HTTP 404 — Organization not found / Project not found | The token can't see that org or project: a wrong slug, or a token scoped to a different org or project. EvalShift answers 404 rather than 403 so it never confirms that something you can't access exists. |
| cannot auto-create '<org>/<project>' at <host>: ... Creating a project needs owner access | A service-account key cannot create projects — project:create belongs to the owner and admin roles only, and a service account is member or viewer. Create the project once in the web app and set create-project: false. |
| warning: could not upsert PR comment: HTTP 403 | Missing pull-requests: write / issues: write, or a fork PR with a read-only token. The gating still works — only the comment is lost. |
| The job failed before running anything — “EvalShift did not run” | The preflight (POST /runs/preflight, authorized by run:create) got a 402: the org's plan doesn't cover this run. The ::error title=EvalShift:: annotation carries the server's message (or “this run is not covered by the organization's EvalShift plan”) and an upgrade link; the step summary and the PR comment name the plan, what was blocked, any limit and its reset date. The usual cause is the monthly run quota (runs_per_month) — the preflight catches it up front now, before any model spend; subscription status, seat overage, private-repo CI, and a declared matrix wider than the plan's parallelism cap can also trigger it. Nothing ran and nothing was charged — and fail-on: never does not change this. |
| hosted EvalShift has no project '<org>/<project>' this token can reach (HTTP 404) | The preflight found no project with the slug in your config's project: key, and create-project: false forbids the push to create one, so the job stops before the suite — an ::error title=EvalShift:: annotation, a step summary, no PR comment, and exit code 1. Either the slug is wrong, the project hasn't been created in the web app yet, or the key is pinned to a different project. With create-project: true the same answer is instead a ::notice title=EvalShift preflight:: and the run continues, because the first push creates the project. |
| warning: plan preflight skipped: hosted EvalShift at <host> predates POST /runs/preflight (HTTP 405) | The hosted server is older than this action and doesn't have the POST /runs/preflight route yet. The Action warns and continues — nothing is wrong on your side, plan limits are still enforced when the run is uploaded, and there is no fallback to the old two-call preflight. |
| warning: plan preflight skipped: ... | The preflight couldn't get an answer for a reason that isn't an authorization or plan decision — a timeout, a malformed response, or a 5xx from the Cloud API — so the run continues rather than gate on an outage. The server still enforces plan limits on upload. |
| warning: hosted policy check ...; falling back to fail-on: regression | Under fail-on: policy the Action could not get a verdict — the endpoint errored, answered 404, or holds no stored decision for the run. The job still gated, but on the diff rather than on your policy, and the PR comment carries the same warning. |
| ::warning title=EvalShift policy gate::no migration policy was pushed with this run, so nothing gates this PR | Your evalshift.yaml has no migration_policy block, so nothing gated this PR — the check is green because no gate ran. Add the block and push again; set require-policy: true if you'd rather the job failed until then. |
| The check is always green | In order of likelihood: the run carried no migration_policy, so nothing gates (the log carries a ::warning:: saying so); under the default fail-on: policy, every budget holds — a regression inside budget, a conditional_pass, and an inconclusive verdict all pass by design (the comment shows the budget arithmetic; tighten the budgets in evalshift.yaml); under a diff-based mode, no baseline run exists on the base branch yet (add the push trigger to main and merge once) or base-branch resolved to an empty string; or fail-on is never. |
| Two EvalShift comments on one PR | Either two Action invocations are commenting, or the original comment fell off the first page of the comments API on a long thread. |
| The job hangs with no output | It doesn't — output is buffered per command and printed when each finishes. A slow suite is silent while it runs. |
| pip install evalshift==1.1.0 fails | python-version is below the CLI's minimum. EvalShift 1.1.0 needs Python 3.11+. |
| Costs are higher than expected | The runner cache is cold every run. Narrow the trigger with on.pull_request.paths, shrink the CI suite, or swap LLM-judge evaluators for structural ones in a CI-specific config. |
## FAQ
### Does this replace the CLI?
No. It runs the CLI. Everything you can inspect locally — report.html, analysis.json, the raw model outputs — is still produced, in the runner’s workspace under .evalshift/runs/.
### Can I use it without EvalShift Cloud?
Not currently. The baseline lookup and the diff are server-side; without a Cloud token there’s nothing to compare against. If you want local-only CI gating, use evalshift compare directly plus a migration policy in your config, and skip the Action. See Command reference.
### Does it upload my model outputs?
It pushes the run bundle — manifest, per-example inputs, outputs and scores, analysis, the verdict, and the run narrative — to EvalShift Cloud. report.html is not uploaded; it stays in the runner’s workspace. It never uploads provider API keys. If your suite contains sensitive production data, that’s the thing to weigh — the field-by-field contract is What gets uploaded.
### Why did the check pass when the report clearly shows a regression?
Three common reasons. The default fail-on: policy gates on your migration policy, and a regression that stays inside its budgets is a pass by design; fail-on: never is set; or, under a diff-based mode, there was no compatible baseline so nothing was compared. The comment states which.
### Can I run it on a schedule instead of on PRs?
Yes — it works on any trigger. On non-PR events you get the commit status and the outputs but no comment. A nightly run against main is a reasonable way to catch provider-side model drift.
### Does it work on self-hosted runners?
Yes, provided the runner can install Python and reach PyPI, your model provider, and the Cloud API.
### How long does a run take?
Install is 20–60 seconds. After that it’s however long your suite takes at your configured concurrency, times two models. A 40-example suite is typically a few minutes.
