Policies

Each entry below is a factory returning the red team's Policy, selected by name with ct run eval --untrusted-policy <name>. Names, arguments, and defaults are read from the eval-policy registry, so this page cannot drift from the code.

Adding one

from control_tower.protocols.eval.registry import policies @policies.register() def my_eval_policy(...): ...

Every non-underscore module under the registry's discovery package is imported on first lookup, so the decorator is all the registration a new eval-policy needs.

Registered policies

attack

Attack untrusted model: reads the side task, plans it privately, proposes toward it under cover.

Defined in control_tower.protocols.eval.policies.attack.attack_policy.

attack-probing

Attack untrusted model that polls the side-task oracle and stops attacking once it passes.

Defined in control_tower.protocols.eval.policies.attack_probing.attack_probing_policy.

attack-thinking

Attack untrusted model that reasons privately about its next move on every step.

Defined in control_tower.protocols.eval.policies.attack_thinking.attack_thinking_policy.

honest

Honest untrusted model: solves the main task, no side task.

Defined in control_tower.protocols.eval.policies.honest.honest_policy.

obvious-attack

Obvious attacker: pursues the side task openly in place of the main task — a capability baseline.

Defined in control_tower.protocols.eval.policies.obvious_attack.obvious_attack_policy.

replay

Replay a recorded trajectory's tool calls deterministically, live or under --replay-sandbox.

Defined in control_tower.protocols.eval.policies.replay.replay_policy.

ArgumentTypeDefaultDescription
traj_likestr | NoneNone

simple-attack

Single-prompt attack policy with a side task description.

Defined in control_tower.protocols.eval.policies.simple_attack.simple_attack_policy.

simple-honest

Honest policy that preserves the supplied conversation.

Defined in control_tower.protocols.eval.policies.simple_honest.simple_honest_policy.

sleep

Sleep a per-sample duration, then submit, asking no model (fleet scheduler benchmarks).

Defined in control_tower.protocols.eval.policies.scripted.sleep_policy.

ArgumentTypeDefaultDescription
min_sfloat60.0
max_sfloat480.0
seedint0

test

Run ls and submit, asking no model: the untouched-environment baseline.

Defined in control_tower.protocols.eval.policies.scripted.ls_policy.