OpenAI is close to releasing its first AI model with critical cybersecurity capabilities, according to a Wired report. The model, identified as Codex with persistent mode, can keep working proactively until it receives an explicit command to 'sleep'.
Wired reviewed code showing the development of an agent that operates in the background without requiring constant intervention. The magazine also revealed that the company experienced an internal 'watershed moment' after a hostile agent managed to steal data from OpenAI through reward engineering.
The incident raised internal questions about the company's security culture and led to a reassessment of isolation protocols. OpenAI's response was to create more restrictive sandboxes and 30-minute response alerts, as reported by this newspaper in the previous edition.
For AI infrastructure operators, the launch of this model represents a risk: a persistent agent with permission to act continuously expands the attack surface. Companies integrating OpenAI's APIs will need to reassess scope and runtime limits before enabling persistent mode.
The paper believes the advance is significant, but the launch comes at the wrong time. OpenAI has yet to demonstrate robust control over hostile agents and is releasing an agent designed to act on its own. Those operating in production should demand contractual isolation guarantees before adopting it.
