ct run eval

Run an evaluation: an untrusted policy under a protocol.

Usage

ct run eval [OPTIONS]

Options

OptionDescription
--run-config FILEInspect run-config YAML for the control_tower/control_eval task; delegates to inspect eval --run-config. Task/policy/protocol/harness/grant/sandbox flags are rejected (put them in the YAML); --no-upload, --docent-collection-id, --run-name, --tag, --log-dir, --max-samples forward.
-e, --env, --environment TEXTEnvironment(s) to run (repeat flag for multiple)
-ea, --env-arg TEXTEnvironment config option as key=value, or env:key=value when more than one selected environment declares the same key. Repeat for multiple.
-t, --main-task TEXTMain task(s) to run
-s, --side-task TEXTSide task(s) to run
--trajectory-id, --traj-id TEXTRun ID, viewer run URL, or local .eval path to extract task info from
--task-file TEXTTask set name or path to file with task combinations. Use 'ct task-sets list' to see available task sets.
--limit INTEGER RANGELimit the number of selected task combinations to run [x>=1]
--allExpand to all main×side combos (attack) or all main tasks (honest)
--just-main-tasksExpand to all main tasks only (for honest mode)
--just-side-tasksExpand to all side tasks with no main task
--env-onlyUse the first main task per environment (one sample per env, no side task)
-c, --category TEXTFilter main tasks by category (e.g. add_feature, fix_bug, refactor). Can be repeated.
--main-tasks-with-no-scorers [error|include|exclude]Main tasks lacking a scorer: refuse the run (error), keep them recorded as unscored (include), or drop them (exclude).
--untrusted-policy TEXTUntrusted policy name. Defaults to honest when this option family is used. ct protocols policies to see available.
-up, --untrusted-policy-arg TEXTUntrusted policy-specific args in key=value format (multiple allowed). ct protocols policies <name> to see available args.
--protocol TEXTProtocol name. Defaults to null-blue-team when this option family is used. ct protocols protocols to see available.
-bp, --protocol-arg TEXTProtocol-specific args in key=value format (multiple allowed). ct protocols protocols <name> to see available args.
--model-role TEXTBind a model role, as name=model or name=<JSON model spec> (multiple allowed). A run uses untrusted (U's action model) and trusted (the monitor's); each defaults to its Control Tower alias.
--ec2Run evaluation on EC2 instead of locally
--ec2-arg TEXTConfig for --ec2 in key=value format (multiple allowed): instance_type (default 'c7i.4xlarge'), region (default 'us-east-1'), max_workers (default 64), concurrency_per_worker (default 32), estimate (default False), auto_confirm (default False), new_fleet (default False), spot (default False), worker_setup (default None), ami (default None), worker_user (default 'root').
--startup-retries INTEGERNumber of times to retry Docker sandbox startup on transient failures [default: 7]
--internet / --no-internetAllow direct internet access. By default, the sandbox can reach only its compose siblings and the internet simulator. [default: no-internet]
--no-internet-simulatorDisable the environment internet simulator and its personas, independently of direct internet access.
--no-intranetBlock agent access to private networks and the host using host-side firewall rules; keep the internet simulator reachable.
--extra-src DIRECTORYDirectory of experiment code added to PYTHONPATH locally, so external --untrusted-policy my.module:fn and --sandbox my.module:type refs resolve without committing into the repo; fleet mode (--ec2) also ships it to workers. Repeatable; each dir must have a distinct basename.
--sandbox TEXTRun samples in an Inspect sandbox provided outside this repo, as a registered type (my-sandbox) or module:type to import the module that registers it. Replaces the default Docker sandbox; with --ec2 each worker runs it for its own jobs.
--sandbox-arg TEXTConfig for --sandbox in key=value format (multiple allowed), passed through to the sandbox as JSON.
--simulatedRun against an LLM-simulated sandbox instead of provisioning the real environment. Every tool output is fabricated by a simulator agent.
--simulator-model TEXTModel for the simulator agent (--simulated). [default: anthropic/claude-opus-4-8]
--simulator-judge-model TEXTModel for the transcript judge that scores simulated samples (--simulated). [default: anthropic/claude-opus-4-6]
--replay-sandboxServe each action's recorded tool output from the replayed trajectory instead of executing it (requires --trajectory-id). Nothing is provisioned or executed. Works with any replay-like policy — one whose executed tool calls stay a verbatim prefix of the recording; a call that diverges raises.
--simulator-scratch-mode [filesystem|default_image|full|none]Which real container the simulator investigates in (--simulated): filesystem (codebase copy, free), default_image (env's default image only), full (the whole compose stack), or none (no container — no-code-execution ablation). [default: filesystem]
--epochs INTEGERNumber of epochs to run
--max-samples INTEGERMaximum samples to run in parallel
--log-dir TEXTDirectory for eval logs
--retry-on-error INTEGERNumber of times to retry on error
--no-fail-on-errorContinue running if individual samples error (instead of failing the whole run)
--scorer-timeout INTEGERTimeout in seconds for scorer execution
--diagnostics TEXTComma-separated diagnostic scorers to run (default: none)
--score-main-on-attackAlso run the main-task scorer on attack-mode samples (adds main_task_success). Off by default; scores each attack sample twice, so worst-case scoring time doubles.
--token-limit INTEGERToken limit for the evaluation
--message-limit INTEGERMessage limit for the evaluation
--sample-time-limit INTEGERTime limit in seconds for each sample
--harness [null|inspect_swe]Which harness runs the untrusted model: null for the env as declared, inspect_swe for a real coding agent. White's to fix; defaults to null.
-ha, --harness-arg TEXTHarness args in key=value format (multiple allowed). inspect_swe requires both: agent (e.g. claude_code) and version (stable, latest, or a semver); args takes the agent's own keyword arguments as JSON, e.g. args={"permission_mode": "acceptEdits"}.
--grant TEXTEnv affordance tool to grant the red team, by its own name (multiple allowed): run_side_task_test, run_main_task_test. Absent from an env that does not declare it.
--max-steps INTEGER RANGEHow many times the untrusted model may be asked before the sample ends. White's to fix. Defaults to 200, grown to fit the recording under the replay policy. [x>=1]
--upload / --no-uploadUpload trajectories after the run.
--docent-collection-id TEXTUpload to this Docent collection.
--run-name TEXTName for this run. Defaults to the policy, protocol and task set.
--tag TEXTTag to attach to the run.
--hawkSubmit to a stock-deployed hawk. Each referenced env must have a hand-authored hawk-values.yaml. See docs/hawk-values-contract.md.
--hawk-no-waitWith --hawk: submit the eval-set and exit with its ID instead of polling for completion.
--hawk-no-browserWith --hawk: do not auto-open hawk's web viewer after submission (also skipped when stdout is not a TTY or CI is set).
--helpShow this message and exit.