Among the numbers published at the GPT-5.5 launch, one explains the pricing decision: the model uses roughly 40% fewer tokens than its predecessor to complete equivalent coding tasks.
That's the figure reconciling a doubled per-token price with the possibility of stable final costs. If each task requires fewer round trips, less reasoning text and fewer discarded attempts, total spend can fall even with a more expensive unit.
The metric shift has practical consequences for comparing suppliers. List price became incomplete information, because two models at the same per-token price can produce very different final costs on the same task.
There's an additional, less discussed effect. Consuming fewer tokens also means responding faster and fitting in a smaller context window, which widens the range of viable applications without changing infrastructure.
For operations of any size, the reliable measure remains the same and takes work: run the real task on the candidates, count total spend until the result is finished, and compare that rather than the price list.
