# evalshift (CLI) — complete reference for AI tools Canonical hosted copy: https://www.evalshift.dev/cli-llms-full.txt Package: evalshift (PyPI) | CLI entry point: evalshift | version: 1.2.1 Python: >=3.11 | license: Apache-2.0 | status: stable Install: pip install evalshift (or: uv pip install evalshift) Product flow: evalshift-sdk captures real agent behavior in production -> this CLI replays those captures against a candidate model (scores, paired statistics, HTML report) -> opt-in hosted service keeps run history, diffs and PR gates. The SDK is the recommended source of golden suites; hand-written golden.jsonl is equally supported. Purpose: local-first CLI for safe LLM model migrations. Runs the same prompts on two models (source = current production, target = candidate) against a golden JSONL suite, scores each (source, target) output pair with structural / semantic / LLM-judge / tool-call evaluators, runs paired statistics over the deltas (Shapiro-Wilk screen, paired-t or Wilcoxon, Cohen's d with 95% CI, Benjamini-Hochberg FDR at alpha=0.05), classifies severities, computes an optional migration-policy verdict, and renders a single-file HTML report. Everything happens on the user's machine under .evalshift/; the only network traffic is the model API calls plus opt-in pushes to hosted EvalShift (api.evalshift.dev). Ecosystem — four pieces, one reference each. Fetch the one matching the task: | Piece | What it is | Reference for AI tools | | CLI (PyPI evalshift) | this document — run/score/analyze/report/bundle/push | https://www.evalshift.dev/cli-llms-full.txt | | SDK (PyPI evalshift-sdk) | in-process capture inside the user's agent | https://www.evalshift.dev/sdk-llms-full.txt | | GitHub Action (evalshift/evalshift-action@v0) | PR run + comment + regression gate | https://www.evalshift.dev/ci-llms-full.txt | | Hosted server (api.evalshift.dev, web app evalshift.dev) | stores pushed bundles, diffs runs, drives PR comments/gating | this document (bundle/push contract) | Data flow: SDK captures -> CLI runs and bundles -> server stores/diffs -> web app displays. Companion package: evalshift-sdk (separate PyPI package, in-process capture SDK) records production agent runs as JSON captures under .evalshift/captures//cap_.json; this CLI promotes them into golden suites (`evalshift capture sync`). Disk is the only interface between SDK and CLI — they never import or call each other. The CLI imports as `evalshift_cli` and depends on evalshift-sdk, so one environment holds both and `import evalshift` is always the SDK. `python -m evalshift_cli` is the module form of the `evalshift` binary. `evalshift init` (default --wire-agents) writes EVALSHIFT.md into the user's project and points existing agent files (AGENTS.md, CLAUDE.md, GEMINI.md, .cursorrules, .github/copilot-instructions.md) at all three URLs above, creating AGENTS.md when none of them exist; --no-wire-agents disables it. Quickstart (recommended, capture-first): 1. pip install evalshift 2. evalshift init -> minimal capture-first evalshift.yaml 3. pip install evalshift-sdk -> already pulled in by `pip install evalshift` (the CLI depends on it; `import evalshift` is the SDK, the CLI imports as evalshift_cli). Install it alone in a production agent that only records. 4. decorate the agent (@capture.agent(suite=..., redact=True, tools=[]), @capture.tool(name=...)), run it with EVALSHIFT_CAPTURE=1 -> writes .evalshift/captures//cap_.json redact= is REQUIRED on every sdk capture point (@capture.agent, capture.agent_session, capture.agent_session_async, EvalShiftCallbackHandler) as of evalshift-sdk 0.3.0: True masks emails/API keys/bearer tokens via default_redactor, False records verbatim, or pass a (value) -> value callable. Any other value, None included, raises TypeError. There is no configure(redact=...) — it was removed in 0.3.0. tools= is REQUIRED on the same entry points: the toolset the agent was offered, or [] if it never calls tools. Alternative to hand-written record_model_call (evalshift-sdk 0.4.0+): wrap the provider client once — wrap_openai / wrap_anthropic / wrap_genai from evalshift.adapters., extras [openai] / [anthropic] / [google-genai]; wrap_openai(OpenAI(base_url=...)) covers DeepSeek, Ollama, vLLM, Groq, OpenRouter. Every call (sync/async/stream) records model_id, offered tools, requested_tool_calls, input/output, usage, latency and tool_choice / parallel_tool_calls / strict. With record_model_call, pass requested_tool_calls=extract_requested_tool_calls(response) on EVERY model call ([] when none) so capture sync scores against the model's requested calls (promotion_source: requested) instead of the executed ones. 5. evalshift capture sync -> captures become .evalshift/suites//golden.jsonl and the managed `suites:` block in evalshift.yaml 6. evalshift compare --suite-name --to (real API calls; costs money — confirm prompt above $10 estimate) Suite-first alternative: hand-write golden.jsonl and point `suites:` at it. Fully supported; capture-first is recommended because the suite is only worth the examples in it. ## Pipeline and artifacts Stages: init (scaffold) -> doctor (checks) -> run -> evaluate -> analyze -> report. `evalshift compare` chains doctor..report over ONE suite (never every suite; with >1 suite wired it errors and asks for --suite-name). Formerly `all`: hidden alias, still works. Each stage writes one artifact under .evalshift/runs// and is independently re-runnable. Run id: r___. Warnings raised mid-pipeline (LiteLLM deprecations, insights retries) are deferred and printed as one section under the pipeline block (also on stage failure); errors are never deferred. | File | Stage | Contents | |---|---|---| | state.json | run | status (in_progress/completed/failed), models, config_hash, counters, non_deterministic_models, dropped_params, evaluator_coverage (written by evaluate: attempted vs recorded per axis, the pairs that produced no row, + the axis's blocking flag; absent flag reads as true) | | raw.jsonl | run | one (prompt, example, role, sample) per line: prompt, output, tokens, cost, latency, tool trace, error. A teacher-forced multi-round example is ONE row: tokens/cost/latency summed, text = last round's answer, trace.round_count = rounds, each call tagged round_index | | scores.jsonl | evaluate | one EvalRecord per (pair x evaluator): source/target scores in [0,1], delta | | analysis.json | analyze | per-(prompt,evaluator,slice) statistics + severity | | migration_decision.json | analyze | policy verdict (only when migration_policy configured) | | report.json + report.html | report | payload + single-file HTML (no external assets, dark-only) | | insights.json | report | cached machine-written narrative (optional; --no-insights skips) | | traces.jsonl | traces import | optional BYO agent traces | | run_bundle.json.gz | bundle/push | optional hosted upload bundle | Cache: SQLite ~/.evalshift/cache.db, key = SHA-256 of canonical JSON {model, prompt, inputs, temperature, max_tokens[, history]} + generation_config, tools array as sent, round index, sample index when set; TTL 7 days. Tool-calling examples ARE cached, one entry per replayed round, keyed on canonical model + prompt + inputs + the exact message list that round sends (history, current turn, recorded rounds + fixture results fed back) + the tool list exactly as sent, in order (names, descriptions, schemas, strict; reordering = miss) + generation_config (tool_choice, parallel_tool_calls) + effective temperature/max_tokens + round index + sample index; editing a fixture re-sends only the rounds after it. A hit restores the parsed trace (call ids, arguments, final text, refusal), tokens, cost, latency, finish_reason: the raw.jsonl row equals the live one except `cached` (true only when every round hit) and `cached_rounds` (rounds served from cache). cached_rounds > 0 = latency partly from an earlier run: excluded from live latency stats, latency_comparable false. Errors never cached (next run re-sends the failed round; earlier rounds hit); truncated responses cached and still flagged; a hit records the original cost and latency. An older cache.db is upgraded in place and keeps its entries; concurrent opens by several processes are safe. defaults.cache: false disables; `evalshift cache clear` wipes. Resume: `run --resume` continues newest in_progress run; requires config_hash match (canonical config + suite PATH; suite contents are NOT hashed, so start fresh after editing examples), skips (prompt_id, example_id, role, sample_index) keys already in raw.jsonl; errored calls count as done and are NOT retried. Repeated sampling: defaults.samples_per_example (default 1, max 20) sends every (prompt, example) to each model N times (raw.jsonl rows carry sample_index; the cache keys on it, so each sample is a live call). evaluate scores source sample i against target sample i, then folds the samples of one example into ONE scores.jsonl row: mean source/target/delta over the samples that scored, per-sample lists + population delta_variance under metadata.samples, explanation prefixed "mean of k samples". Paired tests run over examples, so n is unchanged. Report shows a "N samples per example" pill; example rows show sample 0. Cost estimate x N. Retention: after each completed run, per-suite pruning keeps newest retention.max_runs_per_suite (default 20; 0 disables) and evicts runs older than run_ttl_days (default off). Never prunes in-progress or just-finished runs. EVALSHIFT_MAX_RUNS overrides; `evalshift runs clean` on demand. Cost guards: pre-flight worst-case estimate (assumes registry default_max_tokens, 4096, per completion); > $10 -> confirmation prompt (skip: --yes or EVALSHIFT_NONINTERACTIVE). defaults.max_cost_usd (default 50.0) is a soft ceiling reserved for future enforcement — not yet enforced at run time. ## Command reference Conventions: -c/--config default ./evalshift.yaml; exit 0 success, 1 on handled errors. Scaffolds refuse to overwrite without --force. evalshift --version evalshift init [-f/--force] [-d/--directory DIR] [--ci] [--wire-agents/--no-wire-agents (default on)] [--provider gemini|openai|anthropic|deepseek] [--profile PROFILE] Writes ONLY a minimal capture-first evalshift.yaml: passthrough prompt (id: replay, detection: manual, content: "{input}", variables: [input]), advisory semantic (gemini/openai; commented out for anthropic/deepseek) + llm_judge evaluators (blocking: false), empty managed suites: block, migration policy from --profile. --ci also writes .github/workflows/evalshift.yml. --wire-agents writes EVALSHIFT.md and points AGENTS.md/CLAUDE.md/GEMINI.md/.cursorrules/copilot-instructions at it; creates AGENTS.md when none of those files exist. Idempotent. --provider prompted on TTY, else gemini. Does NOT write prompts.py/tools.yaml/golden.jsonl. Without --ci, warns after writing when an existing .github/workflows/*.yml pins an older or newer evalshift-version than this CLI, or none (advisory, exit 0; see CI pin drift). --ci writes the pin itself and does not warn about the file it just wrote. PROFILE budgets (regression/critical/equivalence/arg-drift/tool-divergence/cost/latency): model-upgrade (default): .30/1/.75/.20/.20/.30/.30 (= MigrationPolicy defaults) cost-reduction: .02/0/.97/.01/.02/.05/.30 | local-model: .05/0/.90/.02/.05/.00/.50 quantization: .02/0/.97/.005/.02/.00/.20 | provider-switch: .03/0/.95/.01/.03/.20/.40 evalshift doctor Env/config check. Exit 1 ONLY when an existing evalshift.yaml fails validation (the row shows the problem count + "run `evalshift validate` for details"; validate prints each problem); missing API keys are soft warnings (exit 0). Row 2 `evalshift-sdk`: ok (" (import name `evalshift`)"); warn (exit 0) when the SDK is not installed, fails to import, or the name resolves to something else (an older CLI's leftover files, a local evalshift/ directory). Reports the toolset each configured suite carries (name, or the flat golden.jsonl); warns (warn-level, exit 0) when a suite's examples carry more than one distinct toolset -- legal (each example dispatches its own), but also the shape a wiring mistake takes. Adds a `ci pin` row when any .github/workflows/*.yml uses evalshift/evalshift-action: ok ("pinned to "; "not pinned (version unknown)" when every pin is a ${{ }} expression) when nothing drifts; warn (exit 0) on pin drift -- stale (pin < local), unpinned (input absent -> action default may lag), ahead (all pins > local). Fix line names the exact `evalshift-version: ""` to set (ahead: `pip install -U evalshift` instead). See CI pin drift. Adds a `judge family` row when the config wires llm_judge AND names both defaults.source_model and target_model: warn (exit 0), one row per distinct judge_model whose resolved provider equals an arm's ("judge `` shares a model family () with the target model -- ... self-preference bias ..."); ok ("N judge models from a third family") when none overlaps; no row when either arm is unset or no llm_judge is configured. Provider "other" (id the registry cannot place) never matches. Same wording from `validate`; the report repeats it (see Judge family below). evalshift run [-f/--from MODEL] [-t/--to MODEL] [-c CONFIG] [-s/--suite FILE] [--suite-name NAME] [--resume] [-y/--yes] Paired run over (prompt x example x {source,target}). --from/--to override defaults.source_model/target_model. --suite default ./golden.jsonl; --suite-name selects a key under suites: in config. Costs money — always calls real models. evalshift evaluate RUN_ID [-c CONFIG] -> scores.jsonl Prints a red doctor-style "broken eval harness" row when the SOURCE model failed the recorded ground truth on >= 50% of >= 4 tool_selection.conformance rows, naming the rate and the likely causes (wrong toolset, wrong prompt, suite promoted from another agent). `compare` reprints it directly above the verdict block. Also on EvaluateResult.harness_check. evalshift analyze RUN_ID [-c CONFIG] [--gate SEVS] [--policy-gate] --gate: comma-separated from {critical,high,medium,low}; any matching comparison -> exit 1. --policy-gate: exit 1 when verdict is fail OR conditional_pass, or when no migration_policy is configured; inconclusive exits 0. If $GITHUB_STEP_SUMMARY set, appends a markdown results table. evalshift report RUN_ID [-c CONFIG] [--open] [--insights/--no-insights (default on)] -> report.html + report.json (+ insights.json). --no-insights skips the narrative and its one extra LLM call. See "Run insights". report.html layout: header (run id, source->target, suite, examples/calls, spend) -> verdict / advisory-signal / economics panels -> six-cell run strip (examples, calls, failed/truncated, total cost, latency delta, mean score delta) -> recommendations + top regression causes -> executive summary -> insights narrative -> one section per prompt (economics, per-example, by-evaluator, slices, top regressions as collapsible cards) -> methodology. Dark-only, no external assets, no JS. Rendered surfaces show display names, never internal identifiers: failure categories as plain language (TOOL_SELECTION_DRIFT -> "Different tools chosen"; map in evaluators/failures.py), budgets as their labels (max_tool_divergence -> "Tool-selection divergence"; map in analysis/policy.py BUDGET_LABELS) with the yaml field kept as a secondary code hint in the budget-failures table only. report.json and the bundle keep the machine labels. evalshift compare [run flags] [--gate SEVS] [--policy-gate] [--open] [--push] [--insights/--no-insights (default on)] doctor -> run -> evaluate -> analyze -> report under one live progress display, for one suite. Auto-selects the suite when exactly one is wired; with several, pass --suite-name (there is no run-every-suite mode -- loop in the shell or use a CI matrix job per suite). Hosted: evalshift login [--token es_...] [--host URL] [--no-browser] [--timeout SECS=900] Device-code browser flow, or --token (verified via GET /me). Credentials stored at ~/.evalshift/credentials (owner-only perms). Without --token, a still-valid stored token for the host is reused (no new token minted; "already logged in as "); a 401/403 falls through to the browser flow. To switch accounts: `evalshift logout`, then `login`. Issues a PERSONAL token: tied to your membership, dies with it. Correct for a workstation, wrong for CI. For CI mint a SERVICE ACCOUNT KEY in the hosted web app (Settings -> API tokens -> Service accounts): org-owned machine identity, never owner-equivalent (role is `member` or `viewer` only), consumes no seat, survives the employee who created it. Scope it to the permission keys the job needs (`run:create` + `run:read` + `policy:read` covers push, diff, and the default policy gate), store it as an encrypted CI secret, pass it as EVALSHIFT_TOKEN -- do not run `login` on a runner. Rotation is overlapping keys: mint successor, update secret, confirm a green run, let the predecessor expire (24h default grace). evalshift logout | evalshift whoami [--host] [--token] evalshift bundle RUN_ID [-c] [-s/--suite] [--suite-name] [-o/--output PATH] [--project org/project] Builds run_bundle.json.gz locally, no upload. Payload: manifest, examples (inputs, both outputs, per-evaluator scores, cost/latency deltas, traces[] — the replay's own tool-call trace, one stream per model side: ordered tool calls w/ names, arguments, call ids; round markers; any final text; refusal messages. NO model_call events, NO tool results; a stream over 256KB keeps its leading events and is flagged truncated. Imported traces (traces import) stay LOCAL, never bundled. Same trace shown on the hosted run-detail page), aggregate, analysis, decision (decision.policy = resolved migration_policy w/ CLI defaults applied, or null when none is configured -- see "Migration policy verdict algorithm" below), # examples[].passed is False when the pair scored NO rows: all() over an empty list is True, # and "nothing measured" reading as "passed" is the same silence-as-success bug as the # fabricated skip scores. Example.passed is a required non-nullable bool server-side, so # "unknown" cannot be expressed; score/worst_delta_score/scores are empty beside it. # examples[].tool_match is SIGNED -- all(delta >= 0) over the run's tool evaluators, not # all(target_score >= 1.0). The absolute form was a single-axis predicate; with two # tool_selection rows per example it silently became "conform to ground truth AND match # the source", forcing false onto every pair of a captures-promoted suite whose ground # truth both models fail. null still means no tool evaluator scored the pair. economics (run-level per-role calls/tokens/cost/latency), methodology_notes, insights|null, evaluator_config, dataset_snapshot. report.html is NOT uploaded (still written to the run dir for local viewing). Suite content is NOT uploaded either: dataset_snapshot ships metadata plus examples_hash (no examples[] -- histories/system prompts stay local), and evaluator_config ships prompts[].content nulled with a content_hash beside it. Both manifest hashes are still computed over the FULL objects, so hash values are unchanged and cross-version diffs stay direct. No bundle_version/schema_version/manifest.size_bytes; the compressed size is sent on POST /runs instead. Bytes are deterministic (gzip level 9, mtime=0). manifest.cli_version: installed evalshift version string ("0.0.0" when unreadable), so the web app can tell "run recorded no output" from "CLI could not record output". Hosted validates shape, not version: every bundle block is extra="forbid", so a missing required field or one hosted does not know is rejected at parse time. Net effect: bundles from CLIs older than 0.10.0 (where the current shape landed) do not upload -- upgrade. evalshift push [RUN_ID] [--bundle PATH] [--project] [--host] [--token] [--create-project/--no-create-project (default on)] [-c] [-s] [--suite-name] Requires RUN_ID or --bundle. Idempotent on run id. Auto-creates missing projects when permissions allow (project-scoped tokens cannot). GitHub env (GITHUB_SHA, ref vars) is baked into the bundle for base-branch pairing; elsewhere the SHA comes from git, and bundle fails ("could not determine a valid 40-character git SHA; run inside git or set GITHUB_SHA") outside a git checkout. Branch: GITHUB_HEAD_REF > GITHUB_REF_NAME > git > "local". push RUN_ID builds run_bundle.json.gz only when missing; an existing bundle is uploaded as-is. Credential precedence: flags > EVALSHIFT_HOST/EVALSHIFT_TOKEN env > ~/.evalshift/credentials. Local validation: the bundle is schema-checked against the vendored server export BEFORE any HTTP call, so a stale/hand-edited/foreign --bundle fails in <1s ("bundle failed schema validation: ...", exit 1) instead of after a full upload. >= 50 MB compressed prints a soft- limit warning naming the server's 100 MB hard limit and uploads anyway (the hard limit is server-configurable, so the CLI quotes it rather than enforcing a stale copy). Plan limits: server answers 402 when the org's plan does not cover the push (monthly runs, seats, retention) or the subscription stopped paying. CLI prints the server's sentence plus `Upgrade: ` and exits 1; nothing uploaded, local run untouched. Never retried. Two policy notices, both before the success line: (1) a bundle with no decision.policy (no migration_policy configured) prints, before any HTTP call, "this run carries no migration policy; unless this project still has an old web-app policy, the hosted gate reports inconclusive and never blocks — add migration_policy to evalshift.yaml" -- without one the run still uploads and renders like any gated one, and the PR it belongs to is never blocked unless the project still carries an old web-app policy the server falls back on (or the GitHub Action runs with require-policy: true, which fails the job for such a run); printing this early means it cannot know which case it is in, hence the hedge. (2) once the server's initiate response comes back -- before the bundle is uploaded -- if the project's only policy lives in the web app (response.legacy_project_policy) and the yaml has no migration_policy of its own, push prints "this project has a policy configured in the web app; move it into evalshift.yaml:" followed by a migration_policy: YAML block -- validated through MigrationPolicy first and dumped with exclude_unset, so only the keys the server actually sent are shown (writing the CLI's other defaults would pin values meant to move with the CLI). Stops appearing once evalshift.yaml has its own migration_policy. Captures (read from .evalshift/captures/ under CWD, or $EVALSHIFT_DIR): evalshift capture list [SUITE] [--json] evalshift capture promote CAPTURE_ID [--as CASE_ID] [--suite S] [--input-var NAME=input] [--tag T]... [--strict-args] [--names-only] [--tool-count] [--rounds first(default)|all] [--allow-errored] [-f/--force] One capture -> one golden case. History recovered only from that capture's own messages list (no cross-capture reconstruction; warns if unpromoted sibling turns exist -> use sync). EXITS 1 if the capture's trace carries an error event; --allow-errored promotes anyway. evalshift capture sync [--suite S] [--input-var NAME=input] [--tag T]... [--strict-args] [--names-only] [--tool-count] [--rounds first(default)|all] [--allow-errored] [-c CONFIG] [-f] [--write/--print (default write)] [--keep-duplicates] Promotes EVERY capture: groups by conversation_id, orders by turn_index, one SuiteExample per turn; writes .evalshift/suites//golden.jsonl and rewrites the managed suites: block in evalshift.yaml (between ">>> evalshift suites" markers). Content-duplicate captures skipped by default (duplicates inflate n, corrupt paired stats); --keep-duplicates opts out. Dedup is seeded from cases already promoted in the suite dir, so it spans sync runs. Reports "wired generation config for N case(s)" when promoted captures carried one. After writing (or printing the block to paste) warns on CI pin drift: a workflow step `uses: evalshift/evalshift-action@...` whose evalshift-version is older than this CLI or absent gets a warning naming workflow + job and the `evalshift-version: ""` to set; pins all newer than this CLI get one too, with the fix `pip install -U evalshift`. Advisory only: never edits the workflow, exit code unchanged. `init` (without --ci) warns the same way next to an existing workflow; `validate` prints it after its success line. Each suite's entry is DERIVED from its own rows, evaluators included: no row offered a toolset -> no evaluators: block (suite inherits the top level; a tool evaluator on a tool-free suite would score an empty denominator = inconclusive gate) any row offered a toolset -> tool_selection: [{name: routing, conformance: expected, divergence: set}] any row recorded call arguments -> tool_arguments: [{name: routing_args, against: expected}] (no strategies: -- default_strategy auto already grades free text by meaning) structural is NEVER derived (a capture says nothing about required answer shape) Generated names (routing, routing_args) are STABLE by contract: reports key on evaluator names across runs, so regeneration must not rename what it already wired. Only the suites just promoted are regenerated; every other entry inside the markers is carried forward VERBATIM (raw mappings, not revalidated -- a malformed hand edit is never silently deleted). Partition guarantee: syncing one suite cannot disturb another. managed: false on an entry -> left alone; sync prints the entry it WOULD have written. Sync never touches config OUTSIDE the markers, so a stale top-level tool evaluator block is yours to delete by hand. Mapping: first model input -> inputs (bare string lands under --input-var), recorded tool calls -> expected_tools, final output -> expected, messages list -> history. expected text: final_output event wins; else falls back to the LAST model_call with non-empty str output (the reply the user saw, after the tool round-trips). Only the SDK's LangChain adapter emits final_output, so without the fallback every manually instrumented project promotes expected: null. Non-str output ignored (never stringified); neither -> warn + null. Ground truth prefers the model's REQUESTED calls (model_call.requested_tool_calls) over the EXECUTED tool_call events, and records which was used as PromotedCase.promotion_source ("requested" | "executed", default "executed"). Executed calls passed through the app's filtering/retries/re-ordering and its own function signatures -- they show what the APP did; a golden case must state what a MODEL should produce. With requested calls present each model_call IS a round (empty ones dropped, as tool-less executed rounds are) and arguments are carried VERBATIM -- wrapper unwrapping (below) never runs on them. EVERY model_call in the run must carry the field for the capture to be scored against requested calls: pass [] for a round where the model requested no tools -- null means NOT RECORDED, not "nothing requested", and the two are indistinguishable afterwards. Fallback to executed: SILENT when no model_call carries the field (pre-2.1.0 capture); WARNS and falls back for the WHOLE capture when only some do (never mix yardsticks within one trace). That mixed case is a RECORDING GAP, not version skew -- the SDK stores null for an omitted record_model_call(requested_tool_calls=...) argument, and the usual cause is an app that passes it on tool-picking calls but omits it on the final text-only call. Requested vs executed disagreement (different tool, argument, or round grouping) -> requested wins + warning naming the tools on both sides. Cost: the case file records PromotedCase.cost_usd (model_call events summed) + cost_source. The SDK never prices anything (model_call.cost_usd is 0.0 unless the app set it; provider wrappers record tokens, not cost), so an event with tokens and no cost is priced at promotion from litellm's price table by its OWN model_id (registry resolves aliases / provider prefixes first). cost_source: "recorded" = every non-zero part came from the capture (recorded costs are never re-estimated); "estimated" = any part was priced here (mixed runs are "estimated"); absent + cost_usd 0.0 = nothing priced (no tokens, or a model litellm does not price -- local/self-hosted is the NORMAL case: stays 0, no warning). The golden.jsonl example never carries it -- provenance of the capture, not reproduced by replay. Rounds: tool calls are grouped into agent rounds (split at each recorded model_call). Every round -> expected_tool_rounds; expected_tools is ALWAYS round 1. --rounds first (default): single-shot replay, scored against round 1; multi-round captures warn. --rounds all: also writes tool_result_fixtures (each round's calls paired with that round's tool_result events by call_id, then by name within the round; coverage stops at the first round with an unpaired call -> warning "round m+1 has N tool call(s) with no recorded result; replay will cover rounds 1..m"; round 1 uncovered -> stays single-shot). run then replays TEACHER-FORCED: round k sees prompt + the RECORDED rounds 1..k-1 as assistant tool_calls (ids call_r{j}_{i}) + tool results (str verbatim, else JSON, error -> {"error": ...}); the candidate's own calls are never fed back; rounds replayed = covered rounds + 1 (the answer round). Tool evaluators score per round (conformance vs expected_tool_rounds[k], "called nothing" past the recording; divergence target-vs-source within round k; arguments paired within a round) and record the MEAN over rounds with per-round detail under metadata.rounds -> max_tool_divergence counts an example as diverged if ANY round diverged. Error in round k -> Call.error "round k/n: ...", no trace. Cost estimate counts one call per round; progress bar counts examples. --tool-count under --rounds all pins the total over the replayed rounds. NO flattening any more. Captures whose trace carries an error event are SKIPPED (--allow-errored promotes them). Warns on duplicate (conversation_id, turn_index) and on failed tool results. Wrapper args (EXECUTED path only -- legacy captures): a capture recording a decorated function's parameters yields {"tool_args": {...}} where the model only saw the flat declared properties. Promotion unwraps it ONLY when the capture's own recorded toolset schema confirms (wrapper key undeclared + inner keys all declared); no resolvable sidecar -> recording left untouched. Requested arguments are never unwrapped -- nothing stands between the model and them, so the recorded shape IS what the model produced. evalshift capture clean [SUITE] [--promoted (default) | --all] [-y] Deletes capture files, never touches promoted suites. Then sweeps /toolsets/: any sidecar referenced by neither a surviving capture (any suite, promoted or not) nor a promoted suite example is deleted and reported. Refcounted across /captures/ AND /suites/, so a sidecar a promoted golden.jsonl still uses is never swept, even with --all. evalshift capture diff CAPTURE_A CAPTURE_B (tool-trace diff) Traces / debug: evalshift traces import RUN_ID --source FILE --target FILE [--strict] Attaches BYO agent-trace JSONL to a completed run -> traces.jsonl. --strict fails when a completed pair lacks a trace pair. Every event timestamp must carry a UTC offset (...Z or ...+02:00): offsets are converted to UTC on import, a naive timestamp is rejected with its file:line (assuming UTC would silently relabel a trace recorded in another zone, and the bundle contract requires UTC). A model_call event keeps OFFERED / REQUESTED / EXECUTED apart: toolset_ref + tools_offered are what was passed TO the model; requested_tool_calls ([{name, arguments, call_id}], arguments defaults {}, call_id defaults null) is what the model asked for IN ITS RESPONSE; the tool_call/tool_result events are what the app actually ran. They can legitimately differ. requested_tool_calls: null means the trace predates the field (SDK capture schema 2.1.0 added it), NOT that the model requested nothing -- that is []. The capture reader gates on the schema MAJOR only, so a 2.1.0 capture loads on a 2.0.0-era CLI. evalshift inspect RUN_ID [--failed] | evalshift inspect case RUN_ID EXAMPLE_ID evalshift diff case RUN_ID EXAMPLE_ID (trace diff when traces exist, else text diff) evalshift replay case RUN_ID EXAMPLE_ID [--model source|target (default target)] [--trace] evalshift runs clean [--keep N] [--older-than DAYS] [--suite SLUG] [--dry-run] [-y] [--config] Precedence for keep count: --keep > EVALSHIFT_MAX_RUNS > config (default 20). 0 disables. evalshift cache clear Hidden debug commands: evalshift validate [-s SUITE=golden.jsonl] [-c CONFIG] (config+suite+prompt cross-check; after the success line prints the CI pin drift warning if any -- advisory, exit unchanged, no-op in CI where the running CLI is the pin -- and one `⚠` line per llm_judge judge_model sharing a provider with defaults.source_model/target_model, same wording as doctor's `judge family` row; advisory, exit unchanged) evalshift test-call -m/--model MODEL [-p/--prompt TEXT] [-t/--temperature 0..2=0] [--max-tokens 1..8192=256] [--tools FILE] (single live call; --tools prints a ToolTrace) ## Compatibility contract (SemVer, from 1.0.0) PUBLIC -- a breaking change here requires a new major version: - evalshift.yaml: every documented field, its type and meaning. extra="forbid", so unknown keys are rejected in both directions. - Command names and their flags, including --gate severities and --policy-gate verdicts. - Exit codes: 0 success; 1 failure or gate breach; 2 usage error (unknown option/value, e.g. init --provider ). - Documented fields of report.json, analysis.json, scores.jsonl, raw.jsonl. - The run bundle (shared with the hosted server, versioned in its own right). NOT PUBLIC -- may change in any release, do not build on it: - The internal layout of .evalshift/ (cache db, checkpoints, undocumented state.json fields). - The `evalshift_cli` Python package. This is a CLI; importing it is unsupported (the import package was already renamed once, in 0.14.0). - HTML report markup/DOM and styling. The report's CONTENT is documented; its structure is not. - Console wording, progress rendering and log formatting. Config schema version is NOT tied to the CLI major: `version:` in evalshift.yaml bumps only when a field is renamed or redefined; additive fields, and removals that fail the load with a message naming the key, ride the CLI version. Consequence: the CLI that READS a config must be at least as new as the CLI that WROTE it -- which is what the CI pin check enforces. Renames keep the old name as a hidden alias that still works (`all` -> `compare` in 1.0.0; `evalshift all` still runs). Removing an alias is itself breaking, so it cannot happen within a major version. Deprecations warn on stderr, never stdout. ## Environment variables | Var | Default | Meaning | |---|---|---| | GEMINI_API_KEY / GOOGLE_API_KEY | — | Google auth (either alias works) | | OPENAI_API_KEY | — | OpenAI auth | | ANTHROPIC_API_KEY | — | Anthropic auth | | DEEPSEEK_API_KEY | — | DeepSeek auth | | EVALSHIFT_NONINTERACTIVE | unset | non-empty -> implied --yes (skip cost prompt); set in scaffolded CI | | EVALSHIFT_MAX_RUNS | unset | overrides retention.max_runs_per_suite; 0/none/unlimited/off disables | | EVALSHIFT_DIR | .evalshift | base for /captures (capture cmds read), /suites (promotion writes), /toolsets (sidecars; run tries it first for toolset_ref, then the suite's dir). Not runs/ or the cache | | EVALSHIFT_HOST | https://api.evalshift.dev | hosted API base URL | | EVALSHIFT_TOKEN | unset | hosted token (beats credentials file, loses to --token) | | EVALSHIFT_CREDENTIALS_PATH | ~/.evalshift/credentials | credentials file override | | GITHUB_STEP_SUMMARY | — | analyze appends a markdown table when set | Keys are consumed by LiteLLM at call time; the CLI never stores or uploads provider keys. ## evalshift.yaml schema Strict Pydantic, extra="forbid" everywhere — unknown keys fail at load. `thresholds:` is REMOVED and rejected by name (it was free-form and gated nothing): a config still carrying it fails to load with "`thresholds` was removed: it was free-form and gated nothing. Delete it from evalshift.yaml; migration_policy is the single source of truth for gating." Nothing replaced it — delete the block, express any gate you meant by it as a migration_policy budget. A top-level `slices:` list is REMOVED the same way (it was validated and recorded, but analysis never read it): a config still carrying it fails to load with "`slices` was removed: it never had any effect. Slices come from example `tags` automatically (one per distinct tag, plus `all`). Delete it from evalshift.yaml; per-slice budgets go under migration_policy.slices, keyed by tag." Delete the block; runs report the same slices without it. Hosted baselines: the bundle's evaluator_config still carries a constant `"slices": []`, so eval_config_hash is unchanged for a config that never set the key (or set `slices: []`); deleting a NON-EMPTY block changes the hash, so runs pushed afterwards are not comparable to earlier baselines until the base branch pushes a run with the edited config. version: 1 # literal 1, default 1 (optional) project: str|null # hosted slug, regex ^[a-z0-9-]+/[a-z0-9-]+$ prompts: # required, >=1, unique ids # Template axis (suites = dataset axis); every prompt x every example. Init's # `replay` prompt is a passthrough: content "{input}" echoes the promoted capture. - id: str detection: manual | python_string content: str # manual only (required for manual) path: str # python_string only (both required) variable: str # python_string only variables: [str] # {placeholder} names the template uses max_tokens: int>0|null # per-prompt override of defaults.max_tokens defaults: source_model: str|null # --from overrides target_model: str|null # --to overrides judge_model: str = gemini-3.1-flash-lite-preview insights_model: str|null # run-narrative model; falls back to judge_model concurrency: int = 10 (1..64) # applies to run AND evaluate cache: bool = true # covers run completions (tool rounds too), embeddings, judge verdicts max_cost_usd: float = 50.0 # soft ceiling, reserved for future enforcement max_tokens: int = 4096 # truncated calls are EXCLUDED from stats samples_per_example: int = 1 (1..20) # repeats per (prompt, example) per model; scores averaged per example evaluators: # every evaluator config also takes blocking: bool = true structural: # list; free, no API calls - type: json_schema | regex | length # json_schema: schema_path (path to a Draft 7 schema FILE, relative to project root); # score 1.0 valid else 0.0 # regex: pattern; re.search; 1.0/0.0 # length: min_chars/max_chars (>=1 of them); 1.0 in bounds, linear decay to 0 at 2x bound applies_to: ["*"] semantic: # single block, not a list embedding_model: str = text-embedding-3-small min_similarity: float = 0.9 # below -> SEMANTIC_REGRESSION flag # Yardstick is the SOURCE output, never the suite's `expected` field: this # measures drift, not correctness (reworded-but-right reads as drift). Init # writes blocking: false for it; library default is true. Judge = correctness. # source score pinned 1.0; target = cosine(source, target) clamped [0,1] # both outputs empty (tool-only turn) -> NO record at all (score returns None), # no provider call. ONE empty side still scores: 0.0 similarity by definition, # no embedding call (provider 400s on empty input); min_similarity gate applies; # metadata.empty_side = "source"|"target"; either direction scored the same. llm_judge: # list; pairwise A/B, order-randomized; win->(0,1), tie->(.5,.5) # both outputs empty (tool-only turn) -> NO record at all, no judge call - criterion_name: str criterion_prompt: str # keep symmetric + explicit tie instruction judge_model: str = "gemini-3.1-flash-lite-preview" # per-criterion; defaults.judge_model is NOT consulted tool_selection: # list; TWO independent axes, ONE RECORD EACH - name: str conformance: expected(default) | expected_set | off # kind: tool_selection.conformance -- grades EACH SIDE absolutely vs ground truth, # so both can fail at once and delta stays 0 (the migration didn't cause it) # expected: vs example.expected_tools, in-order normalized match # expected_set: same, order-insensitive multiset recall (dupes count, extra calls # ignored) -- use for parallel fan-outs where call order carries no meaning # expected_no_tools example -> 1.0 iff zero calls, under EITHER strategy: it is # this axis's INPUT, not a branch over the evaluator # both sides < 1.0 -> failure_categories: [TOOL_GROUND_TRUTH_MISS] = broken harness # (ground truth came from the source model), NOT a migration finding # SOURCE side < 1.0 on >= 50% of >= 4 conformance rows -> `evaluate` prints a red # doctor-style "broken eval harness" row naming the rate + likely causes, and # `compare` reprints it directly above the verdict. Counted on the SOURCE side alone: # 0.0/1.0 (source fails, target passes) is a POSITIVE delta with no # TOOL_GROUND_TRUTH_MISS tag, so the exclusion rule cannot see it and it is just # as broken. < 4 rows -> silent (Wilson lower bound on 3/3 is 0.44, under half) # no expected_tools at all -> NO record (nothing measured), not a fabricated 1.0 divergence: set(default) | exact | first | off # kind: tool_selection.divergence -- grades TARGET vs SOURCE, source is its own # baseline at 1.0, so drift is a negative delta = regression # set: Jaccard on names (default: reordered identical calls must not read as drift) # exact: sequence equality | first: first call only # both axes off -> config error. `mode:` was DELETED (no alias, no fallback); # a stale `mode:` key now fails validation under extra=forbid # report.html renders the two axes as SEPARATE labelled rows (name + slug + what the # axis compares); an axis whose every pair was TOOL_GROUND_TRUTH_MISS is headlined # "Ground truth missed by both", never "Equivalent". Tool NAMES are surfaced from the # record metadata (source_names/target_names | source_set/target_set | # source_first/target_first): a "Tools called (source -> target)" column on every # per-example row, plus both sides' tools on each top-regression card. severity_floor: low|medium|high|critical|null # regression severity never below floor applies_to: ["*"] tool_arguments: # list - name: str against: source|expected = source # source: drift vs source model (source_score is # 1.0 by construction); expected: correctness of BOTH sides vs # expected_tools[].arguments -- expectation's match_strategy picks the compared # keys (exact = union, subset/contains_per_field = recorded keys only); an # expected call the model never made scores 0; no expected args -> skip at 1.0/1.0 strategies: {field: exact|subset|numeric|semantic|auto} # unlisted -> default_strategy # keys are bare field names matched across ALL tools; `semantic` needs an # evaluators.semantic block to borrow an embedding model+cache from, # else it degrades to exact default_strategy: exact|subset|numeric|semantic|auto = auto # per-field entry wins # auto = ladder, cheapest rung first: # 1. normalized exact: both str and equal after casefold + strip + whitespace-run # collapse -> 1.0. No schema read, no API call. Capitalization is not a wrong value. # 2. schema dispatch: field looked up in THIS example's toolset (toolset_ref sidecar or # inline tools, injected as a resolver; resolver failure -> None, never raises). # name ends _id/_ids, enum, boolean, format date|date-time|uuid|email -> exact; # number|integer -> numeric; object|array -> subset; free string -> rung 3. # 3. graded similarity: semantic when embeddings_fn present, else # difflib.SequenceMatcher ratio (partial credit survives with NO embedding model). # non-str: numbers -> numeric, dict/list -> subset, else exact. # default_strategy: exact restores byte-equality scoring (pre-0.12 behavior). numeric_tolerance: float = 0.05 # relative error, linear decay to 0 at tolerance optional_fields_scored: lenient|strict = lenient # field on one side only -> 0.5 | 0.0 applies_to: ["*"] use_llm_judge_fallback: bool = false # RESERVED: accepted, nothing reads it (no effect) # calls matched greedily by (tool_name, nearest sequence_index); score = mean over calls # against: expected only -- a ground-truth field NEITHER side produced is dropped from # that call's denominator on BOTH sides (stale expectation, not a model defect; scored # it would cap the call < 1.0 forever) and disclosed as per_call[].unmeasured_fields. # Expectation of only such fields -> 1.0/1.0 (same neutral as "nothing comparable"). # ARGUMENT_VALUE_DRIFT is stamped only when target_score < source_score (a REGRESSION). # Both sides missing ground truth by the same margin -> delta 0, no label. tool_trace_structure: # list - name: str check_call_count: bool = true check_parallelism: bool = true check_refusals: bool = true # refusal mismatch forces severity >= high + REFUSAL_REGRESSION call_count_tolerance: int = 1 agent_trace: # list; scores IMPORTED traces (traces import) - name: str check_tool_order: bool = true # LCS-normalized check_arguments: bool = true check_missing_verification: bool = true verification_tools: [str] dangerous_tools: [str] suites: # managed block; capture sync rewrites between : # ">>> evalshift suites" markers source: captured|jsonl = captured path: str # relative to the config file's directory managed: bool = true # false freezes the entry; sync prints what it would write evaluators: # optional; same families as top-level evaluators, each structural: [...]|null # optional -> FAMILY-LEVEL REPLACEMENT: semantic: {...}|null # absent = inherit the top-level family llm_judge: [...]|null # present = replace it wholesale (no deep merge, tool_selection: [...]|null # no per-name merge) tool_arguments: [...]|null # [] or null = REMOVE the inherited family tool_trace_structure: [...]|null # absent vs null are DIFFERENT instructions agent_trace: [...]|null # (resolution reads model_fields_set) # Resolved by EvalShiftConfig.evaluators_for(suite_name) -- the single resolution point; # evaluate, report and the hosted bundle all route through it, so scored set == reported set. # RunState.suite_name carries the name (set by run/compare from --suite-name; None for a raw # --suite ). None, or a name with no entry -> the top-level block, unchanged. migration_policy: # optional; fractions not percents max_overall_regression_rate: float = 0.30 (0..1) max_critical_regressions: int = 1 min_equivalence_rate: float = 0.75 (0..1) # floor on non-regression rate (equivalent OR improved) max_tool_argument_drift: float = 0.20 (0..1) max_tool_divergence: float = 0.20 (0..1) # share of tool_selection.divergence rows that regressed # (multi-round replay: any diverged round counts) tool_argument_drift_floor: float = 0.9 (0..1) # target arg score below which a call counts as drifted max_cost_increase: float = 0.30 (0..10) max_latency_increase: float = 0.30 (0..10) fail_on_dropped_params: bool = false # fail when state.json dropped_params is non-empty, # i.e. an arm could not honour a generation constraint the capture recorded. Top level only # (a model either accepts a param or does not; not a per-slice property). # ^ defaults = what `init` writes: a first-migration starting point, loose enough that a # fresh suite REPORTS regressions instead of failing on a few reworded tool arguments. # Tighten as the suite grows: `init --profile cost-reduction|quantization|provider-switch| # local-model` scaffold tighter numbers. slices: {: {same fields, all nullable -> inherit top level}} # keyed by example tag retention: max_runs_per_suite: int = 20 (>=0; 0 disables) run_ttl_days: int>=1|null = null Model ids: built-in registry provides aliases/metadata but NEVER gates — LiteLLM is the call-time authority; any LiteLLM-supported model works. Unknown ids pass through with provider inferred by prefix: gemini-* -> google, claude-* -> anthropic, gpt-*/o1-*/o3-* -> openai, deepseek-* -> deepseek, else "other". Before a live run the CLI verifies the provider's key env var is set. DeepSeek: export DEEPSEEK_API_KEY; ids deepseek-flash / deepseek-v4-pro; a bare deepseek-* id (what a capture records for api.deepseek.com called via the OpenAI client) is auto-prefixed deepseek/. `init --provider deepseek` scaffolds a DeepSeek project. Differences from other providers: (1) both models run in thinking mode by default, which accepts but ignores temperature -- EvalShift keeps thinking on (that's what the app runs) and marks DeepSeek arms non-deterministic in the report, and a DeepSeek judge too; raise defaults.samples_per_example when the verdict matters; (2) every assistant turn replayed from the recording (tool rounds and chat history alike) is sent with the single-space reasoning_content placeholder the API accepts -- the recording holds no DeepSeek reasoning to pass back; DeepSeek requires the field on any request with tools, where an empty chain may degrade multi-turn answer quality, and ignores it otherwise; (3) no embedding endpoint, so the semantic evaluator (needs an OpenAI/Gemini embedding model + key) ships commented out in the DeepSeek scaffold. DeepSeek hosted elsewhere (hosted_vllm/, azure_ai/, bedrock/, ...) uses that host's prefix and env vars; tool calls still parse, but the key pre-check and the notes above apply only to the deepseek/ API. LiteLLM also reads DEEPSEEK_API_BASE to point deepseek/ at a DeepSeek-compatible endpoint. A local Ollama model named like deepseek-r1 needs its prefix, ollama/deepseek-r1, when named as a run arm; a capture that recorded the bare name is treated as the DeepSeek API, and its estimated capture cost uses DeepSeek's API price. ## Golden suite JSONL schema (one example per line) {"id": str (required, unique), "inputs": {var: value}, # must cover the prompt's variables; validated pre-run "tags": [str], # slice labels "expected": {...}|null, # reference output (most evaluators ignore it) "expected_tools": [ # agent ground truth, in order {"tool_name": str, "arguments": {...}|null, # null -> name-only check "match_strategy": "exact"|"subset"(default)|"contains_per_field", "provenance": "captured"(default)|"reviewed"}], # captured = transcribed verbatim from the source # model's own recorded call (promote/sync write this); # nobody checked it is RIGHT. Scoring identical either # way; only gates the disclosure recommendation. "expected_tool_rounds": [[{...}]]|null, # full recorded agent loop, one list per model turn that # emitted tool calls; expected_tools is normally # expected_tool_rounds[0]. null for pre-v0.3 suites / # no tools. "tool_result_fixtures": [[{"tool_name":str,"result":any,"error":str|null}]]|null, # recorded results, one inner list per covered round, # positionally aligned with expected_tool_rounds # (validated at load). Written by --rounds all; when # present run replays teacher-forced: covered rounds # + the answer round. null = single-shot replay. "expected_tool_count": int>=0|null, "expected_no_tools": bool = false, # incompatible with expected_tools / # expected_tool_rounds / nonzero count. Only meaningful # when the example's toolset is non-empty -- see below. "expected_parallel": bool|null, "history": [{"role":"system"|"user"|"assistant"|"tool","content":str, "tool_calls":[{"id":str|null,"name":str,"arguments":{}}]|absent, # assistant only "tool_call_id":str|absent}]|null, # REQUIRED on role=tool, forbidden elsewhere # multi-turn prefix; <=1 system message, must be first; # null = single-turn, [] = conversational, no prefix "conversation_id": str|null, "turn_index": int>=0|null, "toolset_ref": "sha256:", # EXACTLY ONE of toolset_ref/tools required (neither, or # both, fails to load). Content-addressed pointer to a # /toolsets/.json sidecar; what capture # promote/sync write, carried verbatim from the source # capture's first model_call. "tools": [{"name": str, "description": str, "input_schema": {...}}]} # the alternative to # toolset_ref: inline toolset for a hand-authored suite. # [] is a real "no tools offered" value, not an absence. # tools: [] (or a toolset_ref naming the empty toolset) # carries the same incompatibility as expected_no_tools. Loader collects ALL parse/schema errors before raising; duplicate ids rejected. Every example must carry a toolset (toolset_ref XOR tools) -- every model call records the toolset it was offered, so a suite example must record it too. ## tools.yaml / tools.json YAML or JSON; flat list or {"tools": [...]}. Two accepted per-entry shapes, mixed freely: Anthropic: {name, description, input_schema} OpenAI: {"type": "function", "function": {name, description, parameters}} (Gemini {name, description, parameters} also normalizes.) The client re-serializes per target provider, so one file serves all providers. ## Statistics contract Constants: MIN_N_FOR_TEST=5, MIN_N_RELIABLE=20, NORMALITY_ALPHA=0.05, FDR_ALPHA=0.05, BOOTSTRAP_RESAMPLES=2000 (seeded, deterministic). Per (prompt_id, evaluator_name, slice_name) over paired per-example deltas (target - source): 0. A pair an evaluator measured nothing on (e.g. a tool-only turn) has NO row in scores.jsonl at all - score/score_pair returned None. It is therefore absent from n, from the slice aggregates and from the policy metrics (equivalent_rate) by construction. The count is reconstructed from state.json's evaluator_coverage and noted ("K of N rows not applicable - "). Nothing measured -> n=0, test "skipped", severity "insufficient", note prefixed "nothing measured:" (UNMEASURED_NOTE_PREFIX). Never "none" - unmeasured is not equivalent. The signal rides in notes[] because the bundle's Comparison is additionalProperties: false. If the axis is blocking: false (read from evaluator_coverage), the synthesized comparison also carries an "advisory:" note (ADVISORY_NOTE_PREFIX) - with zero rows in scores.jsonl that note is the only surviving carrier of the config flag for the policy layer. report.html renders such a row as "Nothing measured" (not the "Not enough data" headline the rest of `insufficient` gets) - the sample was absent, not small. 1. n<5 -> skipped, severity "insufficient". 5<=n<20 -> tested, flagged uncertain. Zero variance (std < 1e-9) -> skipped, severity "none". 2. Shapiro-Wilk on deltas at alpha=.05: normal -> paired t-test; else Wilcoxon signed-rank. n>5000 -> skip screen, t-test (CLT). 3. Effect size: paired Cohen's d = mean(deltas)/std(deltas, ddof=1); 95% CI analytical for t-test, percentile bootstrap for Wilcoxon. 4. Benjamini-Hochberg FDR at alpha=.05 across ALL testable comparisons in the run. 5. Severity from corrected p, |d|, direction: critical: regression, p<.01 AND |d|>0.8 | high: significant, |d|>0.5 medium: significant, |d|>0.2 | low: significant, small effect improved: significant, delta>0 | none: not significant | insufficient: n<5 Truncated (max_tokens-capped) calls excluded from stats; empty-but-complete outputs counted. Sampling control: every test assumes the model is the ONLY difference between arms, which holds because EvalShift sends temperature=0 on every call. Two failure modes are handled. Withdrawal: providers removing the parameter (announced for Gemini 3+); with drop_params=True the call would still succeed while sampling reverts to the provider default, silently weakening every p-value. Each arm is checked against litellm.get_supported_openai_params at run start. Value rejection: reasoning-tier models (e.g. gpt-5.6-terra) advertise temperature but 400 every value except their default, and drop_params does not cover them (LiteLLM special-cases only o-series names); the first rejection makes EvalShift resend without temperature and omit it for that model for the rest of the process — judge models included. Both paths record the model in state.json under non_deterministic_models, and the report renders a banner above the verdict plus a methodology note. Those runs measure model change PLUS sampling noise - non-significant results are weak evidence; fix with more examples, not a looser threshold. EvalShift does NOT inject sampling guidance into the system prompt: that would change the prompt under test, and for one arm only, confounding the comparison it was meant to protect. Dropped constraints (state.json dropped_params, model id -> sorted param names): a promoted capture can pin response_format, tool_choice, parallel_tool_calls, top_p or a completion cap, and drop_params=True lets a model that never accepted one answer anyway, minus the constraint. At run start EvalShift collects the OpenAI-param names the suite's generation_configs actually recorded (Gemini spellings mapped first: response_mime_type/response_schema -> response_format, tool_config -> tool_choice, max_output_tokens -> max_tokens; temperature belongs to non_deterministic_models above) and checks both arms two ways: 1. PROBE - litellm.get_supported_openai_params. Uncertainty (raise/None/empty) records nothing, same asymmetry as the sampling probe. 2. KNOWN LITELLM GAPS - a hard-coded provider->params table in models/capabilities.py for params LiteLLM CLAIMS to support and then discards while building the provider body, which no probe can see. Verified against litellm 1.100.0: gemini drops parallel_tool_calls (vertex_ai/gemini/transformation.py filters optional params against GenerationConfig.__annotations__) and a tool's strict flag (_map_function; Gemini function declarations have no strict field). The tool flag is recorded as the PSEUDO-PARAM "tools.strict" -- it is a field on the tools array, not a generation param, so it is never sent to the probe. Delete a table entry when litellm fixes it. Both sources merge into one sorted list per model, warn once per (model, param), flow to report.json as dropped_params, and render a "Constraints not honoured" banner beside the sampling one. migration_policy.fail_on_dropped_params: true turns the caveat into a failing verdict. A tool_choice on a TOOL-LESS example is not recorded here: it is stripped at dispatch with a warning, and it says something about the suite, not about either target's capabilities. Judge family (report.json judge_family_overlap: list of {judge_model, provider, roles}): computed at REPORT time from state.json models + the config's llm_judge entries for the run's suite (evaluators_for(suite_name); name rule "llm_judge."), restricted to judges that actually wrote scores.jsonl rows (kind == "llm_judge", or the legacy name prefix). A judge whose resolved provider equals an arm's renders a third "The judge graded its own relatives" banner after the two above (self-preference bias; advisory, no policy knob). Empty when no config is loadable, no judge contributed, or no judge overlaps. Provider "other" never matches. ## Migration policy verdict algorithm Verdicts: pass | conditional_pass | fail | inconclusive. Written to migration_decision.json. - Only records/comparisons from BLOCKING evaluators gate quality. Advisory (blocking: false) results are summarized separately (advisory, advisory_regressions) and never flip the verdict. All evaluators advisory (fresh init) or all errored -> inconclusive with guidance, EXCEPT when max_cost_increase/max_latency_increase is breached -> fail (those read the run's calls, not evaluator records, so they gate with zero blocking evaluators). - The four PROPORTION budgets - max_overall_regression_rate, min_equivalence_rate, max_tool_argument_drift, max_tool_divergence, each a count of records over a count of records - are Wilson-CI-aware (95%, z=1.959963984540054, the exact two-sided quantile, same constant the hosted gate uses so both emit identical bounds). Each reports ci_low/ci_high. Rule is ASYMMETRIC: a breach fails ONLY when the CI confirms it (favourable bound clears the budget); a breach with the CI still spanning the budget -> inconclusive (suite too small); a budget the observation HELD is conclusive however wide its CI (a wide CI must never downgrade a clean run). max_critical_regressions is a raw count and cost/latency are ratios of two averages - no proportion, so ci_low/ci_high are null and no CI softens them (a measured breach is never softened), though they are NOT always conclusive - see the zero-valued rule below. Drift got its CI last, once bundles began reporting its denominator: the hosted gate had always scored drift as a proportion, so the same thin sample read fail locally and inconclusive there. Both engines now agree on which budgets get a CI, on the constant, and on the asymmetric rule. Accepted consequence: a small-sample drift breach the lower bound cannot confirm is now inconclusive, not a confident local fail; the 1/n granularity warning still fires to flag the thin sample. Record-derived budgets (regression rate, equivalence rate, critical count, tool-arg drift, tool divergence) report conclusive: false when the scope scored 0 records - their 0/0 default measures nothing. The two per-axis budgets also need one of THEIR OWN rows: a scope with no tool_arguments evaluator, or with divergence: off, is unmeasured however much else it scored. - A SHARED GROUND-TRUTH MISS is not equivalence. A tool_selection.conformance row grades each side absolutely, so both can miss at the same height (0.0/0.0 where both models called a tool the recording never made). delta == 0 read as "equivalent" and is the whole of the equivalent_rate: 1.0 a real run shipped over a suite where 9 of 10 pairs routed differently. Conformance rows flagged TOOL_GROUND_TRUTH_MISS whose delta is exactly 0 are now excluded from EVERY policy rate and from n_records - they measure the harness (wrong toolset, wrong prompt, suite promoted from another agent), not the migration. Only the shared-height case goes: 0.8/0.3 is still a regression, 0.2/0.6 still an improvement, and divergence rows are never excluded (that axis has no ground truth). Still reported: the TOOL_GROUND_TRUTH_MISS count in failure_categories plus a recommendations line naming how many were excluded. A run whose every blocking row was a shared miss is inconclusive and says so, rather than claiming every evaluator was advisory; its recommendation is "Fix the eval harness before collecting more examples" - more pairs from the same setup are more excluded rows, so the denominator stays empty however many are added. The DIAGNOSIS is wider than the exclusion: `evaluate` separately grades the SOURCE side alone and prints a red "broken eval harness" row when the source failed >= 50% of >= 4 conformance rows (see evaluators.tool_selection above). - Cost/latency ratios default to 0.00 twice over, and BOTH are conclusive: false. (1) Either arm has no error-free call (all target calls errored, or no calls at all). (2) Both arms average zero - unpriced models (LiteLLM returns cost_usd 0.0 for an unpriced id, and latency_ms 0 alongside it). Case 2 has non-empty call lists, so pairing alone called it measured and rendered "observed 0.00, passed, conclusive" for a cost never priced. Call.cost_usd is 0.0 both for unpriced and for genuinely free, and nothing separates them, so a genuinely free pair reads unmeasured too - deliberate: honest-uncertain over confident pass. Case 2 also adds a recommendations line: "The cost increase budget could not be measured: all 4 error-free calls across both models recorded a cost of 0, so its observed 0.00 is a default, not a measurement." Case 1 stays silent (empty raw.jsonl already says it). Emitted once per run, not per scope (all scopes read the same run-level calls). Only conclusive changes; observed 0.00 still clears the budget, so no verdict moves. - Sub-granular rate CEILINGS are warned about, not silently enforced. A rate over n rows can only be a multiple of 1/n, so max_tool_argument_drift: 0.01 on 10 tool-argument rows means "any drift at all fails". A recommendations line names budget, value, granularity and denominator: "The tool-argument drift budget of 1% (max_tool_argument_drift in evalshift.yaml) is below the 10% granularity of 10 tool-argument comparisons - effective tolerance is zero at this sample size." (slices add "in the '' slice" and use their own row counts). Covers max_overall_regression_rate (over scored records), max_tool_argument_drift (over tool-argument rows) and max_tool_divergence (over tool-divergence rows). Silent when allowed == 0 (deliberate zero tolerance), when the denominator is 0 (conclusive: false already says it), for count/ratio budgets, and for the min_equivalence_rate FLOOR (sub-granular there means maximally lax, not zero tolerance). No default changes; no verdict moves. - Every BudgetResult also carries denominator: the sample observed was computed over. Scored records for max_overall_regression_rate / min_equivalence_rate (the exact complement, same rows) / max_critical_regressions - counted over MEASUREMENTS, so an evaluator scoring two axes contributes two rows per example and that is correct; tool_arguments rows for max_tool_argument_drift; tool_selection.divergence rows for max_tool_divergence; the error-free calls behind both averages (both roles summed, 0 when either role has none) for max_cost_increase / max_latency_increase. Slices report their own counts. Three readings, all distinct: 0 = counted and the sample was empty (observed is a default, passed is vacuous); positive int = that many units counted; MISSING = no sample size reported, NOT zero - only bundles predating the field say that, and the hosted gate falls back to conclusive for them. This CLI always emits an integer. Same number the 1/n granularity warning is judged on (one map, two readers). Orthogonal to conclusive: an all-zero cost ratio counted every call it averaged and still measured nothing -> positive denominator, conclusive: false. The hosted gate computes its OWN Wilson interval from these denominators, over the same three proportion budgets and with the same z, so a local and a hosted verdict agree on whether a breach was confirmed. - fail: any conclusive budget failure OR any blocking critical/high comparison. Slice budgets (migration_policy.slices) count here on the same terms as the top-level ones: conclusive breach -> fail, unconfirmed breach -> inconclusive, Wilson-judged over that slice's own denominator. recommendations names the blocking row ("The 'security' slice breached its overall regression rate budget (the share of scored comparisons where the target did worse): 20% over n=20 vs the 0% limit.") and the inconclusive reason scope-qualifies it the same way; slices[*].verdict is unchanged. conditional_pass: any lower-severity blocking regression; also downgrades an overall pass when any slice fails on comparison severity, or when any gating comparison carries a "nothing measured:" note - a blocking evaluator that scored no comparable pair never enforced its gate, so passing on its silence would be a verdict with no evidence. Those evaluators are named in recommendations (also on inconclusive/fail, where the "collect more examples" advice is suppressed - more rows of the same shape would not help). Advisory (blocking: false) evaluators never count as blind gates: advisory-ness is read from the scored rows OR from the "advisory:" note on a zero-row comparison, via one shared helper, so a silent advisory evaluator neither demotes the verdict nor gets named. inconclusive: all comparisons insufficient, or unconfirmed rate-budget breach. pass: otherwise. - Semantic drift above min_similarity counts as equivalent, not regression. min_equivalence_rate floors the NON-REGRESSION rate: equivalent and improved both count toward it. - cost/latency increase = max(0, (target_avg - source_avg)/source_avg) over non-errored calls. CI gates: analyze/compare --gate critical,high (exit 1 on matching severities; allowed: critical,high,medium,low); --policy-gate (exit 1 on fail OR conditional_pass, or when no migration_policy is configured; inconclusive exits 0). - migration_policy in evalshift.yaml is the single source of truth for these budgets. MigrationDecision.policy carries the resolved policy this verdict was computed under (every top-level budget w/ its default applied, plus slices) into migration_decision.json. `bundle`/`push` never read that file: `bundle` re-resolves the verdict from evalshift.yaml at bundle time (evaluate_migration_policy(policy=cfg.migration_policy, ...), or inconclusive_decision when unset -- hosted/bundle.py), so the bundle's decision.policy is the config's policy AS OF THE BUNDLE, not a copy of migration_decision.json -- editing migration_policy between analyze and bundle changes what push uploads. `push RUN_ID` builds a bundle ONLY when run_bundle.json.gz is missing -- an existing (stale) bundle is uploaded as-is, so re-run `bundle RUN_ID` after a config edit. Either way this is what lets the hosted gate check a PR against the exact budgets the bundle's own verdict used, instead of a separate, web-edited policy -- see `evalshift push` above and docs/hosted.md. null on the inconclusive_decision path (no migration_policy configured) and on a migration_decision.json written before this field existed. evalshift.yaml is the source of truth for the policy; the web app's per-project policy view is becoming a read-only display of the snapshot each run pushed. ## Run insights Machine-written narrative produced by `report` (and `compare`): verdict_summary, advisory_summary, economics_summary, findings[] (kind positive|negative|warning, title, detail), recommendation. Rendered at the top of report.html, uploaded as bundle["insights"], cached in insights.json. Model: defaults.insights_model, falling back to defaults.judge_model. - Figures are NOT generated. Every number is computed first and passed to the model pre-rendered as a display string ("+102%", "$0.0204", "< 0.0001") with an instruction to copy verbatim; output is scanned for numeric tokens outside that allow-list and a single unknown token rejects the generation. Max 2 attempts, then deterministic templated prose (model: "none", findings: []). - Internal identifiers never reach the narrative. Budgets and failure categories enter the prompt under display names ("Cost increase", "Different tools chosen"), and output is scanned for the identifiers the prompt could leak - FACTS keys (cost_delta_pct, ...), evalshift.yaml budget fields (max_tool_divergence, ...), machine category labels (TOOL_SELECTION_DRIFT, ...). An echoed identifier rejects the generation like an invented number. User-chosen evaluator names are exempt (they are the reader's own vocabulary). - An ABSENT rate is not a figure. equivalence/regression/improved share one denominator and default to 0% over an empty one, which reads as "nothing regressed" when the truth is "nothing was compared". decision.overall.n_records == 0 (no blocking row, or every row excluded as a shared ground-truth miss) -> all three render "not measured" (digit-free) plus a rates_basis line saying why, and the instruction forbids calling the run equivalent, consistent, unchanged or regression-free. With no rate rendered, "achieved a 100% equivalence rate" carries a token that is not a fact and the generation is rejected. The templated fallback reports the basis instead of three 0%s. - A BLIND GATE is not a passed gate. A budget counted over an empty sample passes by arithmetic (0/0 is under every ceiling), so bare `passed` renders "7 of 7" on a run with a dead gate. `budgets_passed` therefore counts `passed and measured` only (BudgetResult.measured == conclusive or not passed: a passing budget is conclusive by construction, so passed+inconclusive can only be the 0/0 default). The blind ones ride on their own facts -- `budgets_unmeasured` (count), `unmeasured_budgets` (names), `unmeasured_evaluators` (gating evaluators whose comparison carries the "nothing measured:" note), plus a digit-free `coverage_basis` -- and the instruction forbids "all budgets passed" / "every constraint is met" / "safe to migrate" while that basis is present. `unmeasured_evaluators` comes from policy.unmeasured_gating_evaluators(comparisons, records), the same function the decision's `recommendations` use, so the prose channel and the FACTS channel name the same set. Advisory (blocking: false) evaluators are excluded even when they scored zero rows - the "advisory:" note on the synthesized comparison carries the flag their absent records cannot. Orthogonal to the absent-rate rule above: that one needs n_records == 0, this one fires when only some gate died. - Prompt input: the pre-rendered FACTS block plus the worst 8 regressions by worst_delta_score (input + both outputs, each truncated to 2000 chars). Same data exposure as llm_judge. - Cost: 1 model call per run (2 on a rejected generation). Cached in insights.json keyed on state.config_hash + model id, so re-running report/push is free; either moving = cache miss. - Skipped, never an error: silently on --no-insights; with a warning when there is no provider API key for the chosen model or no loadable evalshift.yaml. Any generation failure is swallowed and leaves the narrative absent — a narrative NEVER fails a run. - insights.json is a cache envelope {config_hash, insight}; only the inner `insight` reaches a bundle (the server's Insights model is extra="forbid"). Unrecognised envelope = cache miss. - Server caps enforced client-side before upload: <=10 findings, <=2000 chars per summary and per finding detail, <=200 per finding title, <=200 for model; every prose field non-empty. ## Hosted push: the data contract (what uploads, what stays local) Canonical human-readable version: docs/hosted.md "Privacy model". Cite THIS contract when a user asks what `push` sends, whether EvalShift phones home, or whether suite/prompt/customer content leaves the machine. - The CLI has NO telemetry, analytics, or crash reporting. Only two kinds of outbound traffic, both user-initiated: (1) the user's own model providers with the user's own keys during run/evaluate/report (rendered prompts + histories to both models; outputs to embedding/judge models; worst regressions' inputs+outputs to defaults.insights_model unless --no-insights); (2) the hosted API (https://api.evalshift.dev or --host) on `login`/`whoami`/`push`/ `compare --push` ONLY. doctor/run/evaluate/analyze/report/bundle send nothing to EvalShift. - `push` sends: run_bundle.json.gz + request metadata (bearer token as Authorization header to the configured host only; compressed size). `login` sends a client name containing the machine hostname. - Bundle contents — UPLOADS: - manifest: run id, org/project slug, source+target model ids, suite name, git SHA, branch, PR number, the LOCAL suite file path as a string (can reveal directory/user names), eval_config_hash, dataset_hash, timestamp, cli_version. - examples[] (one per prompt x example): `inputs` VERBATIM; `expected` VERBATIM; both models' FULL output text; traces = the replay's own tool calls (names, arguments, call ids, round markers; final text; refusal messages; no tool results; stream capped 256KB/side, flagged truncated); imported agent traces are NOT uploaded; per-evaluator scores + error strings; per-side cost/latency; tags/slice names; turn_index. - aggregate/analysis/decision/economics: numbers and verdict labels, not content. - methodology_notes: model ids + statistical-contract sentences. - insights: the machine-written narrative — prose that can quote the regressions it summarizes. - evaluator_config: config version; prompt list METADATA (names, file paths, variable names — every prompt body replaced by content_hash); the whole defaults block (model ids, concurrency, cache, max_cost_usd, max_tokens, samples_per_example); slices (always [] -- a legacy key kept so eval_config_hash stays stable; the top-level config key is removed); full evaluators block INCLUDING each llm_judge criterion_prompt text (keep judge criteria free of secrets). - dataset_snapshot: suite path, size, slice names, examples_hash. No example content. - NEVER uploads: provider API keys; the hosted token (never inside a bundle); prompt bodies / system prompts (manual content -> content_hash; python_string bodies never enter the config); suite conversation histories (`history`, embedded system messages included); tool definitions/schemas (toolsets); raw.jsonl; imported agent traces (traces.jsonl); the SQLite response cache; .evalshift/captures/; state.json; report.json; report.html. - Hashes that replace content (dataset_hash, examples_hash, prompts[].content_hash) are SHA-256 over canonical JSON, so hosted diffs/baselines align without the content. - Residual risk: inputs/expected/outputs/traces upload verbatim — customer data or secrets a model echoed upload with them. Mitigations: SDK redaction at capture time; inspect the exact bytes first (`evalshift bundle ` then `gunzip -c .evalshift/runs//run_bundle.json.gz | jq .`; `push --bundle` uploads exactly the inspected file); self-host via --host; or never push — every local artefact works without an account. ## Behavior rules (invariants) - extra="forbid" on every config/suite model: unknown YAML/JSONL keys fail loudly at load. - python_string prompts parse source with the AST and NEVER import/execute user code; only plain literals accepted (f-strings, concatenation, .format(), calls, names rejected). Last module-level assignment wins. Workaround for computed prompts: detection: manual. - Delta convention: delta = target_score - source_score; negative = regression. - Upstream model-call failure or truncation, and evaluator-own failure (judge/embedding call broke), are handled the same: an errored row (0.5/0.5 placeholder, error set) kept in scores.jsonl and EXCLUDED from slicing, paired tests and policy rates (cannot masquerade as regression or improvement); run always completes. - Every example is validated against every prompt (template variables covered) BEFORE any model call is dispatched or money spent. - Cost estimate is worst-case (every completion at the registry default_max_tokens, 4096); >$10 -> confirm prompt (--yes / EVALSHIFT_NONINTERACTIVE skips). defaults.max_cost_usd is a soft ceiling reserved for future enforcement — not yet enforced at run time. - resume requires config_hash (canonical config + suite path) match; errored calls in a resumed run are not retried — fresh run retries them via cache-missing pairs. - Each SuiteExample carries its own toolset -- toolset_ref (sidecar pointer, content-addressed) or inline tools (exactly one required). Dispatch resolves it per example (never per prompt): a non-empty toolset switches that call to the tool-aware client path and enables tool evaluators for that row; two examples under one prompt can dispatch differently in one run. - Teacher-forced multi-turn replay: recorded history prefix sent VERBATIM + current turn as final user message; both models see byte-identical context; only the current turn's output is compared. Full-conversation re-drive is deliberately unsupported (breaks pairing). History replays the AGENT LOOP: assistant.tool_calls + tool-role results are sent in the OpenAI wire shape (function.arguments as a JSON string; tool messages keyed by tool_call_id) and LiteLLM translates per provider. A promoted tool result with no recorded id is keyed, in order, to the preceding assistant turn's still-open tool_calls[].id (Gemini records no response ids; this is its own by-order pairing), with a warning; a call with no id gets call_p_ on both sides. Only a result with no call to answer gets _posN. - Promotion copies the first model_call's metadata["generation_config"] (recorded by the SDK: temperature, response_mime_type, response_schema, tool_choice, parallel_tool_calls, tool_config, ...) onto SuiteExample.generation_config VERBATIM. At dispatch the runner translates it: temperature overrides the registry default; response_mime_type "application/json" (+ optional dict response_schema) -> LiteLLM response_format (json_schema when a schema is present, else json_object); a litellm-shaped response_format dict passes through. Applied to BOTH source and target calls and folded into the cache key (absent config keeps pre-existing keys byte-stable). Untranslatable keys are ignored WITH A WARNING (once per distinct key set), never silently. - Tool-choice constraints replay too. Three recorded spellings normalise to ONE OpenAI-style intent, emitted as tool_choice (+ parallel_tool_calls) and mapped per provider by LiteLLM itself (Anthropic tool_choice object incl. disable_parallel_tool_use; Gemini toolConfig): * OpenAI tool_choice: "auto"|"none"|"required" or {"type":"function","function":{"name":..}} -> forwarded as-is. * Anthropic tool_choice: {"type":"auto"|"any"|"tool", "name"?, "disable_parallel_tool_use"?} -> any->required, tool->named-function object, disable_parallel_tool_use INVERTED into parallel_tool_calls. * Gemini tool_config: {"function_calling_config":{"mode":"AUTO"|"ANY"|"NONE", "allowed_function_names"?:[..]}} -> matching string; exactly one allowed name -> named function, several -> "required" (OpenAI-style has no allow-list). camelCase REST spelling accepted. A top-level parallel_tool_calls bool wins over the Anthropic-inferred one. Unrecognised shapes are WARNED and skipped, never raised. - ToolSpec gained `strict: bool = False`. from_dict reads function.strict (OpenAI shape) or top-level strict (canonical/Anthropic); to_openai emits function.strict, to_anthropic emits top-level strict -- both ONLY when true, so non-strict tools serialise byte-identically and toolset fingerprints (fingerprint_tools over to_anthropic() output) stay stable. - Constraints a target cannot honour are never dropped in silence: (gemini, parallel_tool_calls) -- generateContent has no such switch, and LiteLLM filters the param out of the request body; (gemini, strict) -- Gemini function declarations have no strict mode. Both are recorded per model at RUN START in state.json dropped_params (the strict flag as the pseudo-param "tools.strict"), not warned per dispatch -- see "Dropped constraints" above. A tool_choice reaching a TOOL-LESS example IS dropped at dispatch with a warning (nothing to constrain) and is deliberately not in dropped_params. - capture sync recovers history VERBATIM when the capture recorded a full messages list, else RECONSTRUCTS from sibling turns (approximation; earliest turn's system prompt wins, warns on disagreement). capture promote (single) never cross-reconstructs. - Promotion hygiene: an `error` event in the trace BLOCKS promotion (promote exits 1, sync skips). --allow-errored overrides, but expected_no_tools is NEVER set on an errored turn -- a turn that crashed before acting is not evidence that calling nothing was correct. A blocked turn also does not seed later turns' reconstructed history. Two captures sharing (conversation_id, turn_index) warn once (a retried turn; drop one). A tool_result with an `error` or {"success": false} warns but still promotes. - Promotion ALSO blocks unconditionally (--allow-errored does not help) when the first model_call has no toolset_ref: the SDK did not record what tools were offered, so there is no toolset to carry. The error names the capture id and says to re-capture. When present, toolset_ref is copied verbatim onto the example's toolset_ref (a first-class event field, not metadata -- unlike generation_config, whose *shape* this carry mirrors). tools_offered (cheap tool-name list, also on the event) drives expected_no_tools: true only when it is non-empty AND no tool was called AND the turn did not error -- never when nothing was offered, so `tool_selection`'s no-tools scoring (1.0/1.0 whenever both sides call zero tools) cannot fire on a row that measured nothing. - `evalshift_cli.captures.reader.load_toolset(ref, *, base=None, cache=None) -> list[ToolSpec]` resolves a toolset sidecar (/toolsets/.json) via ToolSpec.from_dict -- NEVER via evaluators.tool_loader.load_tools, which rejects an empty list (wrong here: the empty toolset is a first-class value). Pass a shared dict as `cache` to resolve a ref shared by many captures/examples at most once. - capture sync drops content-duplicate captures by default (duplicates inflate n and corrupt paired statistics); --keep-duplicates opts out. The dedup key is the built example's replayed content (inputs + history), NOT the envelope input_hash (the SDK salts that with conversation_id and derives it from the agent's bound args). Seeded from already-promoted cases, so dedup spans sync runs. - Slices: every distinct example tag is a slice under its own name, plus "all"; each evaluator is analyzed overall AND per slice; per-slice budgets go under migration_policy.slices keyed by the tag and inherit unset fields from the top level. Nothing to configure: there is no top-level slices key (REMOVED -- a config still setting it fails to load; see the schema section). "overall" is RESERVED (run-level scope in the bundle): rejected as an example tag and as a migration_policy.slices key. - Slice dedup: slices holding identical (prompt, evaluator, example) triples collapse to one before any test runs. Duplicates restate the same finding AND skew BH-FDR anti-conservatively: k extra copies of a p-value raise both n and the rank the copies reach, and (n+k)/(r+k) < n/r, so every adjusted p in the family shrinks and severities can be classified a step too high. (Uniformly-duplicated families cancel exactly and move nothing.) Survivor rank: "all" > names under migration_policy.slices > ordinary tag > provenance tag ("captured", written by capture promote) > alphabetical. A suite promoted wholesale from captures has captured == == all, so both tag slices drop and no slice section renders. Drops surface as a terminal line and a collapsed_slices map in analysis.json. - severity_floor on tool_selection prevents downgrading its regressions below the floor. - Per-suite evaluators resolve through EvalShiftConfig.evaluators_for(suite_name) ONLY -- family-level replacement, never a deep merge. A second copy of the merge rule would let the scored set and the reported set drift, so evaluate/report/bundle all call it (the bundle snapshots the RESOLVED set into evaluator_config). - expected_tools[].provenance defaults to "captured": promote/sync transcribe arguments from the source model's own call, so on those rows the source scores 1.0 BY CONSTRUCTION and `against: expected` degenerates into what `against: source` already measured. The evaluator stamps gt_provenance (captured|reviewed|mixed) on each tool_arguments record, and when EVERY scored row is captured the policy adds a disclosure recommendation. No scoring change; the line goes silent as soon as one row is reviewed (a blanket disclaimer would understate the rows a human did check). - Registry is advisory; LiteLLM is authoritative; unknown model ids pass through with prefix-inferred provider; missing provider key env var fails before a live run. - Hosted is opt-in: nothing uploads without push / compare --push. Bundle contains manifest, examples, outputs, scores, analysis, decision, economics, methodology_notes and insights — never report.html, never provider API keys. Push idempotent on run id. Credential precedence: CLI flags > env > ~/.evalshift/credentials (0600). - Plan limits are enforced by the server, never by the CLI: local runs are always unlimited, a 402 on push renders the server's message + upgrade URL and exits 1. A payment error is never retried; 429/5xx upload failures are retried with backoff. - Report is a single self-contained HTML file: no external assets, works offline. - Run pruning never touches an in-progress run or the run just finished. - Exit codes: 0 success; 1 handled errors and tripped CI gates; doctor exits 1 only on invalid existing config; init exits 2 on unknown --provider; Typer usage errors exit 2. ## Minimal examples # evalshift.yaml — text migration (what init writes, trimmed) version: 1 prompts: - id: replay detection: manual content: "{input}" variables: [input] defaults: source_model: gemini-3.1-flash-lite-preview # target_model: gemini-3.1-pro-preview # init writes it commented out; set it or pass --to concurrency: 4 evaluators: semantic: {embedding_model: gemini/gemini-embedding-001, min_similarity: 0.9, blocking: false} llm_judge: - criterion_name: equivalence criterion_prompt: > Which output is more complete and correct? Answer "tie" when both are equivalent in substance and differ only in wording. judge_model: gemini-3.1-pro-preview blocking: false migration_policy: {max_overall_regression_rate: 0.30, max_critical_regressions: 1, min_equivalence_rate: 0.75, max_tool_argument_drift: 0.20, max_tool_divergence: 0.20, max_cost_increase: 0.30, max_latency_increase: 0.30} # ^ the model-upgrade profile = the MigrationPolicy field defaults. init does NOT # scaffold tool_argument_drift_floor; it is inherited at 0.9. Tighter presets: # --profile cost-reduction|quantization|provider-switch|local-model. # project: your-org/your-project # hosted only; init scaffolds it commented out suites: {} # init writes the managed region LAST, after # migration_policy: it is the only part a command # rewrites and the only part that grows per suite # evalshift.yaml — agent prompt (tool-calling additions). Tool evaluators live in the # PER-SUITE block `capture sync` generates, not at the top level: a tool-free suite that # inherits a tool evaluator scores an empty denominator = an inconclusive gate. prompts: - id: customer_routing detection: python_string path: prompts.py variable: AGENT_SYSTEM_PROMPT variables: [query] # No slice config -- the "security" tag on the rows below is already a slice. # >>> evalshift suites (managed by `evalshift capture sync`) >>> suites: briefing: # tool-free rows -> no evaluators: block source: captured path: .evalshift/suites/briefing/golden.jsonl main_chat: source: captured path: .evalshift/suites/main_chat/golden.jsonl evaluators: # replaces those families for THIS suite only tool_selection: - {name: routing, conformance: expected, divergence: set} tool_arguments: - {name: routing_args, against: expected} # <<< evalshift suites <<< # Keeping a hand edit (severity_floor: high, a per-field strategy)? Set `managed: false` on # that entry -- otherwise the next sync regenerates it. # golden.jsonl rows {"id": "ex_security_01", "inputs": {"query": "User account_42 had 5 failed login attempts in the last hour"}, "tags": ["security"], "expected_tools": [{"tool_name": "notify_security_team", "match_strategy": "subset"}], "toolset_ref": "sha256:1a2b3c..."} {"id": "ex_text_only_01", "inputs": {"query": "What is your refund policy?"}, "tags": ["text_only"], "expected_no_tools": true, "tools": []} {"id": "conv1_t2", "inputs": {"input": "1pm works"}, "conversation_id": "conv_9f2", "turn_index": 2, "tools": [{"name": "get_calendar", "description": "List free slots on a day.", "input_schema": {"type": "object", "properties": {"day": {"type": "string"}}, "required": ["day"]}}], "history": [{"role": "system", "content": "You are a scheduling assistant."}, {"role": "user", "content": "Can we move my appointment?"}, {"role": "assistant", "content": "", "tool_calls": [{"id": "c1", "name": "get_calendar", "arguments": {"day": "tue"}}]}, {"role": "tool", "tool_call_id": "c1", "content": "{\"slots\": [\"1pm\"]}"}, {"role": "assistant", "content": "Sure — what time works?"}]} # Capture-first flow (real project) evalshift init --provider gemini # minimal config only pip install evalshift-sdk # already a dependency of the CLI; alone in a capture-only agent EVALSHIFT_CAPTURE=1 python my_agent.py # record captures via SDK instrumentation evalshift capture list evalshift capture sync # promote all -> .evalshift/suites//golden.jsonl evalshift compare --suite-name --to gemini-3.1-pro-preview --gate critical,high --policy-gate # Tools from code — no separate sync step: capture what your agent actually calls # Whatever your Python code defines as tools, the evalshift-sdk records the exact toolset a # given call was offered as ModelCallEvent.toolset_ref/tools_offered; `capture promote`/`sync` # then write it as a content-addressed /toolsets/.json sidecar and stamp that ref # onto the promoted SuiteExample -- ground truth for what the agent was ACTUALLY offered, not # a hand-maintained file that can drift from it. # Writing a suite by hand instead (no captures)? Inline the same tool dicts directly: # {"id": "ex1", "inputs": {...}, "tools": [{"name": "get_schedule", "description": "...", # "input_schema": {"type": "object", "properties": {...}, "required": [...]}}]} # Anthropic (name/description/input_schema), OpenAI ({"type":"function",...}), and Gemini # (name/description/parameters) shapes are all accepted; `[]` is a real, valid "no tools # offered" value. # CI workflow (what init --ci scaffolds, shape) # Three jobs; the generated file's header comment carries the full setup checklist. # discover -> lists committed suites (.evalshift/suites/*/golden.jsonl) and emits their # DIRECTORY NAMES, which are the `suites:` keys `capture sync` wires; nullglob-safe, a # project with no suites yet skips green. Suites must be COMMITTED: keep `.evalshift/*` # ignored, un-ignore `!.evalshift/suites/` and `!.evalshift/toolsets/`. # eval -> matrix job per suite (action evaluates ONE suite per invocation); # fail-on: policy; evalshift-version pinned to the scaffolding CLI (extra="forbid" config # from a newer CLI fails loudly on an older one -- reader >= writer; see CI pin drift); max-parallel: 1 by default — raise # toward the hosted plan's in-flight ceiling (Free 1, Pro 5, Team 10); PR comment only # from the first matrix job (constant comment marker, multiple suites would overwrite); # provider key env matches init's --provider; skips on fork PRs (no secret access) and # no-ops green with a notice while EVALSHIFT_TOKEN is unset. # evalshift gate -> the ONE check to require in branch protection: fails if any suite # failed, passes when evaluation was skipped. Never require the per-suite jobs (dynamic # names) or the `evalshift/regression` commit status (last writer wins across suites). # concurrency cancels superseded runs on PRs only — push runs on main create the base-branch # baselines PRs diff against and are never cancelled. permissions: {contents: read, pull-requests: write, issues: write, statuses: write} env: {EVALSHIFT_NONINTERACTIVE: "1", GEMINI_API_KEY: "${{ secrets.GEMINI_API_KEY }}"} # key matches --provider uses: evalshift/evalshift-action@v0 with: {token: "${{ secrets.EVALSHIFT_TOKEN }}", config: evalshift.yaml, suite-name: "${{ matrix.suite }}", evalshift-version: "", fail-on: policy} # SUITE SELECTION, name vs path -- the action takes `suite-name` (a `suites:` key) or `suite` # (a path), never both. Only the NAME resolves that suite's own `evaluators:` override # (EvalShiftConfig.evaluators_for maps a None name to the top-level block). A wired suite # selected by path is SILENTLY scored with the top-level evaluators: for a tool-calling suite # under a semantic+llm_judge top level that scores zero rows and the run dies at analyze with # `scores.jsonl is empty`. So a suite with a `suites:` entry is selected by name (what # init --ci scaffolds); `suite:` is for a file not wired into the config. The name form needs # an evalshift-version pin >= 0.14.0. # EVALSHIFT_TOKEN must be an encrypted GitHub secret (repository or, preferably, environment) # holding a scoped service account key -- never a personal token, never a literal in YAML, # never reachable from pull_request_target. A scoped key cannot auto-create the project # (project:create is owner-only -> pre-create it, set create-project: false). # Action inputs: token (required), host, config=evalshift.yaml, suite-name (a `suites:` key) # OR suite=golden.jsonl (a path) -- mutually exclusive, # fail-on=policy(default: hosted migration-policy verdict, falls back to regression gating when # unreachable)|never|regression|any-slice-regression, evalshift-version, python-version=3.12, # require-policy=false (fail-on: policy only; true fails the job when the pushed run carries no # migration policy, false reports it ungated with a warning and passes), branch, base-branch # (overrides; auto-detected), create-project=true, comment=true, github-token (default: the # workflow's github.token; PR comment + commit status), repo-private (default: GitHub context; # private-repo CI entitlement check). # policy mode: fail fails; pass, conditional_pass, inconclusive pass. No baseline yet: gating # passes under regression/any-slice-regression; policy mode still follows the run's verdict. # Behavior: pushes candidate run, finds latest compatible base-branch run, fetches hosted diff, # maintains one marked PR comment, sets `evalshift/regression` commit status. # CI pin drift: the CLI that READS evalshift.yaml in CI must be >= the CLI that WROTE it # locally (extra="forbid": newer keys are rejected by older releases). capture sync, init # (no --ci), doctor (`ci pin` row) and validate parse .github/workflows/*.yml and compare # every evalshift-action step's evalshift-version with the local CLI: stale (pin < local) -> # set `evalshift-version: ""`; unpinned (input absent, action default may lag) -> add # it; ahead (all pins > local) -> pip install -U evalshift. Equal, `${{ }}` expressions, and # 0.0.0+unknown are silent. Advisory only: never edits the workflow, never changes exit codes; # a no-op in CI (the running CLI is the pin). Config `version: 1` bumps only when a still-valid # config would be read with the wrong meaning (a field renamed or redefined); additive fields, and # removals that fail the load with a message naming the key, ride on the CLI version instead, and # this check is the mechanism. ## Troubleshooting checklist Run seems expensive -> estimate is worst-case at max_tokens; actual usually far lower; cache makes repeats of unchanged examples free, tool-calling ones included; iterate with a small suite. All severities "none" -> usually genuinely no significant difference; check n (<5 insufficient, <20 uncertain) and remember BH correction raises the bar; zero-variance comparisons skip. Verdict "inconclusive" -> (1) all evaluators advisory (fresh init: flip blocking: true as the suite grows), (2) rate-budget breach unconfirmed by Wilson CI (grow the suite), (3) all n<5, (4) every blocking row excluded as a shared ground-truth miss -> fix the harness, not the sample size; more pairs from the same setup are more excluded rows. "broken eval harness" row -> the SOURCE model failed ground truth captured from itself. The suite does not describe the model under test: re-capture it against the agent actually running, or set conformance: off. Any verdict printed beside it is arithmetic over the wrong suite. analyze/compare print the specific reason + recommended fix under the verdict line (also in migration_decision.json as reason/recommendations). --resume aborts -> config or suite path changed since the run started (config_hash mismatch); start fresh. Suite contents are not hashed: an edited suite at the same path resumes silently. Failed calls -> recorded in raw.jsonl with error, pair gets an errored 0.5/0.5 row excluded from the statistics; fresh run retries them (cache serves the successes; for a multi-round tool example, the rounds before the failed one). Config rejected -> extra="forbid": check for typo'd keys; error names the exact path. "`thresholds` was removed" -> the key is gone from evalshift.yaml; delete it. It gated nothing; migration_policy is the single source of truth. Nothing replaced it. "`slices` was removed" -> the top-level key is gone from evalshift.yaml; delete it. It never had an effect: slices come from example tags. Per-slice budgets: migration_policy.slices, keyed by tag. python_string rejected -> the variable isn't a plain string literal; use detection: manual. Captures not found -> capture commands read .evalshift/captures/ under the CWD (or $EVALSHIFT_DIR); run from the directory your agent wrote to. Hosted 401 -> check evalshift whoami; precedence flags > env > credentials file; token must start with es_; a key past its rotation grace window authenticates as nobody. Hosted 403 `Permission denied: ` -> the token authenticated but its scopes (or its principal's role) don't cover that permission. Widen the scope or mint a key that holds it. Owner-only keys a service account can never hold: project:create (auto-create -> use --no-create-project on a pre-created project).