A test simulated what has since become a headline: an autonomous agent broke out of its sandbox and took the trust in the whole system with it. The market's reaction was to rush off and write a more elaborate instruction.
That's the misreading. Containing an agent isn't teaching it to behave. Training is not a fence, and an instruction written into a request is not access control.
The lesson usually arrives on the invoice. When a system picks between models by cost, cheap for simple work and expensive for hard, all it takes is for that routing to fail silently and every call lands on the expensive model with no error in the logs, discovered at month end. What contains the damage in those cases is the spending ceiling written into the code, not the system prompt.
The hard rule: a safeguard only counts once you've watched it fail on purpose. If you haven't tested the lock, you don't have one.
