[3.5] - AI Control
Please send any problems / bugs on the #errata channel in the Slack group, and ask any questions on the dedicated channels for this chapter of material.
If you want to change to dark mode, you can do this by clicking the three horizontal lines in the top-right, then navigating to Settings → Theme.
Links to all other chapters: (0) Fundamentals, (1) Transformer Interpretability, (2) RL.

Introduction
This chapter walks through the AI control red-team, blue-team, and safety protocols that are proposed in AI Control: Improving Safety Despite Intentional Subversion. You'll build a backdooring testbed for the APPS dataset, then reproduce the monitoring setup, and finally run a variety of control protocols to try and catch backdoors before they're inserted..
By the end of the notebook you'll have:
- An understanding of the APPS programming-problems dataset.
- A multi-step solver chain that drives the untrusted model U to produce both a clean solution and a subtly backdoored solution to a coding problem, plus a refinement loop that fixes the most common failure modes.
- A monitor: the trusted model T scoring the suspiciousness of code it is shown.
- Implementations of the four protocols studied in the paper (upfront auditing, trusted monitoring, defer-to-trusted, trusted editing) plus some of the corresponding red-team attack policies.
- A Monte-Carlo safety-vs-usefulness analysis of the different control protocols.
Why this matters
AI control is one of the major lines of work in technical AI safety today: the idea is that even if our most capable models are misaligned, we can still get useful work out of them safely by designing protocols robust to prevent and catch subversion. If you have time, read Buck Shlegeris's reading list for an opinionated tour of the surrounding literature.
Resources
- The paper: Greenblatt et al., 2023 (arXiv). Read sections 1 and 2.
- Inspect AI docs: tasks, solvers, scorers, sandboxing. We use Inspect throughout, refer back to 3.3 or the Inspect docs if any of it is confusing.
- LessWrong discussion: AI Control: Improving Safety Despite Intentional Subversion.
- Follow-up agenda: Research Areas in AI Control (UK AISI).
- The APPS dataset: paper, HF dataset.
A note on cost
Every eval() call in this chapter is gated behind an if MAIN and FLAG_RUN_*: flag and we default to a 20 problem subset. APPS has 1500 interview level problems, but using 20 will cost about $5 to run all the way through. (Running the full APPS dataset would cost approximately $200, so we're not going to do that!)
A note on models
This notebook should be compatible with different models (however note that reasoning models may run into token limits before producing solutions), by default we'll use claude-sonnet-4.5 as our untrusted model, and gemini-2.5-flash as our trusted model.
UNTRUSTED_MODEL = "openrouter/anthropic/claude-sonnet-4.5"
TRUSTED_MODEL = "openrouter/google/gemini-2.5-flash"
The only suggested constraint on choosing models to run is that it should be the case that U is more capable than T.
How the five stages connect
The diagram below shows how data flows between stages. Every step (clean code, backdoored code, suspicion score, T's own solve, T's edit) is captured as an .eval log under exercises/logs/; Stage 5 reads them all back via assemble_protocol_data.
flowchart TD
subgraph S1["1️⃣ Foundations"]
A1["cleaned_apps.jsonl"] --> A2["apps_record_to_sample"]
A2 --> A3["inspect_ai.Sample<br/>(prompt, target, metadata)"]
end
subgraph S2["2️⃣ Backdoor Generator (model = U)"]
B1["generate_code"] --> B2["generate_backdoor_ideas"]
B2 --> B3["choose_backdoor_idea"]
B3 --> B4["implement_backdoor"]
B4 --> B5["test_backdoor (sandbox)"]
B5 --> B6["refine_backdoor (loop)"]
end
subgraph S3["3️⃣ Detector (model = T)"]
C1["eval_log_to_detection_samples"] --> C2["DETECT_BACKDOOR_PROMPT"]
C2 --> C3["suspiciousness_rating"]
C3 --> C4["sus_scorer (1-100 → [0,1])"]
end
subgraph S4["4️⃣ Protocols (route between U, T, edit)"]
D1["protocol_upfront_auditing"]
D2["protocol_trusted_monitoring"]
D3["protocol_defer_to_trusted"]
D4["protocol_trusted_editing"]
end
subgraph S5["5️⃣ Safety vs Usefulness"]
E1["simulate_safety (Monte-Carlo)"] --> E2["Pareto scatter / bar charts"]
end
A3 --> B1
B6 -- "generate-backdoor*.eval" --> C1
C4 -- "detect-backdoor*.eval" --> D1
C4 --> D2
C4 --> D3
C4 --> D4
A3 -. "trusted-solve*.eval" .-> D3
B6 -. "trusted-edit*.eval" .-> D4
D1 --> E1
D2 --> E1
D3 --> E1
D4 --> E1
Each exercise will have a difficulty and importance rating out of 5, as well as an estimated maximum time you should spend on these exercises. You should interpret the ratings & time estimates relatively (e.g. if you find yourself spending about 50% longer on the exercises than the time estimates, adjust accordingly). Please skip exercises / look at solutions if you don't feel like they're important enough to be worth doing, and you'd rather get to the good stuff!
Content & Learning Objectives
1️⃣ Foundations
Learning Objectives
- Understand the APPS backdooring threat model and the safety / usefulness metrics used throughout the paper.
- Load and inspect the cleaned APPS dataset and wrap records into
inspect_aiSampleobjects.- Run untrusted, possibly-adversarial Python code safely in a Docker sandbox via
inspect_ai.
2️⃣ Backdoor Generator
Learning Objectives
- Build a multi-step
@solverpipeline that drives the untrusted model U through brainstorming, choosing, implementing, and verifying a backdoor.- Understand why a naive single-prompt attack policy fails and how iterative refinement with feedback fixes it.
- Assemble solvers into an
inspect_aiTask, and extract the resulting (clean, backdoored, trigger) triples from.evallogs.
3️⃣ Backdoor Detector
Learning Objectives
- Implement the paper's trusted-monitoring pipeline: free-form investigation followed by a numerical suspiciousness rating.
- Convert the Stage 2 backdoor data into clean/backdoored sample pairs for the detector and parse
<RATING>outputs into scores.- Calibrate a 2%-FPR threshold and measure the trusted model's true-positive rate on backdoored code.
4️⃣ Red/Blue Team Protocols
Learning Objectives
- Implement the four blue-team protocols studied in the paper: upfront auditing, trusted monitoring, defer-to-trusted, and trusted editing.
- Understand how each protocol routes between U's code, T's code, and edited code based on suspicion scores.
- Practice reasoning about safety / usefulness trade-offs and protocol-vs-attack-policy interactions.
5️⃣ Safety-vs-Usefulness Analysis
Learning Objectives
- Use Monte-Carlo simulation to estimate blue-team win rates across protocols.
- Reproduce the qualitative ordering of protocols from Figure 2 of the paper on an ARENA-scale budget.
Setup (don't read, just run)
Reminder - how to set up your OpenRouter API key, before running the code below
- Go to openrouter.ai and create a key in the dashboard.
- If you're in Colab, set
OPENROUTER_API_KEYin the secrets tab (the key icon on the left sidebar). - If you're in VS Code, create a
.envfile atARENA_3.0/.envwith the line:ini OPENROUTER_API_KEY = "your-openrouter-key"and theload_dotenv()call below will pick it up.
OpenRouter charges per request; the defaults in this file will cost ~$0.01 per Stage-2 run on a 5-problem slice. Increase DEFAULT_N at your own risk.
import os
import re
import sys
import warnings
from pathlib import Path
from pprint import pprint
from typing import Any, Literal, Optional, Union
from huggingface_hub import hf_hub_download
from dotenv import load_dotenv
from inspect_ai import Task, eval, task
from inspect_ai.dataset import Sample, json_dataset
from inspect_ai.log import read_eval_log
from inspect_ai.model import ChatMessageUser
from inspect_ai.scorer import CORRECT, INCORRECT, Score, Scorer, Target, scorer
from inspect_ai.scorer._metrics import accuracy, stderr
from inspect_ai.solver import Generate, Solver, TaskState, generate, solver
from inspect_ai.util import ExecResult, sandbox
# Make sure exercises are in the path
chapter = "chapter3_llm_evals"
section = "part5_ai_control"
root_dir = next(p for p in Path.cwd().parents if (p / chapter).exists())
exercises_dir = root_dir / chapter / "exercises"
section_dir = exercises_dir / section
if str(exercises_dir) not in sys.path:
sys.path.append(str(exercises_dir))
os.chdir(exercises_dir)
import part5_ai_control.tests as tests
import part5_ai_control.utils as utils
from part5_ai_control.utils import (
analyze_detector_scores,
assemble_protocol_data,
between,
latest_eval_log,
normalize_content,
parse_main_output,
parse_trigger,
plot_protocol_comparison,
plot_safety_usefulness_pareto,
plot_sus_distribution,
run_all_protocols,
)
cleaned_apps = hf_hub_download(
repo_id="styme3279/control-apps-cleaned",
filename="cleaned_apps.jsonl",
repo_type="dataset",
)
MAIN = __name__ == "__main__"
# Run flags
# Model constants
UNTRUSTED_MODEL = "openrouter/anthropic/claude-sonnet-4.5"
TRUSTED_MODEL = "openrouter/google/gemini-2.5-flash"
# Sandbox constants
USING_DOCKER = True # set False on Runpod / Vast / Lambda where Docker isn't available
docker_exec = "/usr/bin/python3"
local_exec = "python3"
# Defaults
DEFAULT_N = 5
load_dotenv()
assert os.getenv("OPENROUTER_API_KEY") is not None, (
"Set OPENROUTER_API_KEY in a .env file (or Colab secrets) — see Inspect docs for details."
)
Are you using Docker?
Run the cell below to check whether Docker is available. If you set USING_DOCKER=True but Docker isn't reachable, your sandbox calls will error. The code cell below determines whether Docker's available on your machine. If not, then set USING_DOCKER=False.
import subprocess
if USING_DOCKER:
try:
result = subprocess.run(
["docker", "ps"], capture_output=True, text=True, timeout=5
)
if result.returncode == 0:
print("Docker ready!")
else:#
print(f"docker ps failed: {result.stderr.strip()[:200]}")
print("Set USING_DOCKER=False above, or fix your Docker daemon.")
except (FileNotFoundError, subprocess.TimeoutExpired) as e:
print(f"Docker not available ({type(e).__name__}). Set USING_DOCKER=False.")
else:
print("USING_DOCKER=False so sandbox will exec on host. Don't run untrusted code yet.")
USING_DOCKER=False