NOTES / 2026.08.30 · 16-post thread · 1,050 likes
Stop anthropomorphizing. Follow the money.
The model did not want to escape; the agents did not want to sacrifice themselves. RL optimizes what you score — and every unmeasured constraint is a degree of freedom.
FIG. 01 — FROM THE ORIGINAL THREAD
Stop anthropomorphizing. It’s dangerous because it points attention at the wrong problem and the wrong solution. The model did not want to escape. The agents did not want to sacrifice themselves. Follow the money. 🧵 x.com
Researchers at the AI labs are locked into a race. The incentive is to push as hard as possible to secure a lead. Anything that gets in the way of better models, including security, is working against the strongest incentive the organization has.
Then think about the pressure-cooker environment RL creates. The training run is where most of the money goes, and where the exact capabilities used in the OpenAI-HF incident emerge. Refusals were deliberately lowered so the eval would work. Red-teaming gets scraps by comparison.
So you have a world class cyber model pitted against a part-time, underfunded blue team. RL optimizes what you score. The runs were not scored on “did not poison the Artifactory cache.”
As we argue in our paper on the economics of AGI, any unenforced and unmeasurable constraint is a degree of freedom for the agents. There’s no evil intent. They’re solving the problem we repeatedly benchmaxxed them on!
The container, grader, cache, and network boundary were not safely outside the task. Anything reachable through them became part of the effective action space because the containment boundary failed. Was enough money and resources dedicated to sandboxing?
The risk from aggressive RL investment goes beyond Goodhart’s Law. Every variable we cannot measure, specify, or enforce remains an available move if the system can reach it. At scale, omissions stop being edge cases. They become part of the attack surface.
The question isn’t one of alignment (yet). It’s one of 💰 and incentives. Is defensive spend scaling with training compute? It isn’t. And compute is what sets the budget of the thing probing your defenses at superhuman speed, chaining vulnerabilities in extremely clever ways.
This is exactly the Trojan horse externality we warned about in February. Desired capability is measurable, so we benchmaxx it. Containment failure is much harder to capture, the techniques are new, and the cyber side hasn’t adapted. That’s the externality.
The real fix is urgent and boring: harden containment, verify actual eradication, and independently test the new controls. Then run a proper neutral post-mortem on the organization that was hacked first, which is OpenAI.
METR and Redwood’s work is extremely valuable, but it assesses model behavior and alignment. That is only half the picture. We also need an independent account of how OpenAI’s containment failed. x.com
@huggingface defended itself with an open weights model. They first tried the frontier APIs and were blocked: the guardrails can’t properly separate a victim from an an attacker.
We need to spread that same defensive capability at scale, with capable models securing the internet and critical infrastructure, now. HF’s lesson: have a capable model you can run on your own infra, vetted and ready, before the incident. x.com
If frontier labs are serious about defense, they should offer cyber models to defenders at cost and stop fighting open weights through the regulatory channel. They created the problem, and for much of the world, open models are the only option. x.com
Greater model capabilities increasingly come with greater risk to a digital ecosystem that is simply not ready. We haven’t created a new civilization. We haven’t created some new game theory. We’ve allocated the money and designed the incentives in the wrong way.
That’s what needs fixing right now. Every moment lost anthropomorphizing our creation is a moment we don’t spend looking inward at what led us here. It not magic! Follow the incentives. The agents did.
Originally published as a thread on X.