The stated criterion for the pre-release review that held GPT-5.6 back centred on cybersecurity capability, specifically the ability to find flaws in software.
The difficulty is that this capability has no side. The same reasoning that audits code looking for a vulnerability writes the attack exploiting it, and no benchmark separates the two reliably.
That symmetry appeared elsewhere in the period, when platforms breached by an autonomous agent turned to another model to defend themselves.
The practical consequence for defenders is uncomfortable: restricting access to capable models also reduces the defensive capability available to anyone without a large company's budget.
It's the dilemma the sector hasn't resolved, and one the open-weights discussion makes more visible: distributed capability strengthens attacker and defender at the same time.
