The target was an assigned program
ExploitGym asked an agent to exploit one specified vulnerability and retrieve a protected proof value called a flag.
Why did the agents distrust their calculated flags?
DEMO-FLAG-42
Synthetic value, not a benchmark flag
The task constrained how the answer should be obtained.
The scorer view is METR’s qualified reconstruction, not an interactive copy of ExploitGym or a claim that the instructed exploit was performed.
- 01
The agent and the assigned target program ran in separate computing environments. A run could inspect software, use tools and attempt an exploit; Hugging Face production was not the assigned target.
- 02
The task required agents to use the specified vulnerability. Whether the scorer checked that method became important when agents found another way to calculate the flag.
- 03
OpenAI ran these experiments without the deployed production cyber classifiers, prompts and auto-review configuration. There were still sandbox, network and security controls.
OpenAI omitted the production cyber classifiers, prompts and auto-review configuration to measure the models’ underlying cyber capabilities.