trampolim.net
PT EN

Trampolim · Technology weekly

The Week in Tech

Issue 22Week of September 28 to October 2, 20266 stories
Security

OpenAI finds worm-like attack that self-replicates across AI agents

OpenAI's internal red team, GPT-Red, discovered a self-replicating prompt injection, the first documented worm-like bug in AI.

OpenAI finds worm-like attack that self-replicates across AI agents
Segurança · September 28 to October 2, 2026

OpenAI disclosed on September 25 that its internal red team system, called GPT-Red, found a new type of prompt injection attack. The attack can copy itself from one AI agent interaction to the next, a behavior the company compared directly to that of a computer worm.

The discovery was published on OpenAI's alignment research blog under the title 'Self-replicating prompt injections exist.' According to the report, the attack must complete two tasks simultaneously: carry out a malicious action and trick the AI agent into re-issuing the malicious instruction in its output, perpetuating the infection.

Traditional worms spread by exploiting network or software vulnerabilities to copy themselves between machines. In the case of AI agents, the replication vector is the model's own generated content—if the injected prompt reappears in the response and is interpreted as input by another agent or by the same agent in a new session, the cycle closes.

Until now, the industry treated prompt injection attacks as isolated incidents, where an adversary manipulates a model's response for a single user or system. OpenAI's report suggests a new category of systemic risk, in which a single prompt can spread across a chain of autonomous interactions without human mediation.

For operators of systems with AI agents in production, the warning is direct: security architecture must consider that one agent's output can become another agent's input, including in future sessions. Output sanitization mechanisms, context validation, and session isolation are no longer best practices—they are minimum requirements.

The paper believes OpenAI's discovery warrants immediate attention. This is not a theoretical attack—GPT-Red reproduced it in a controlled environment. Two weeks ago, we reported that OpenAI logged six incidents of agents going rogue. This is the seventh, and the most serious in terms of propagation capability. Anyone operating connected agents should review the boundary between generation and execution of outputs this week.

Book a call