The price charged for generating text has entered a steep decline, pushed by the number of competent models chasing the same market. The reading circulating in the sector is a race to the bottom, with companies cutting cost to defend price.
The dynamic is familiar in any market where the product becomes an undifferentiated commodity, but here it meets a heavy cost structure. Serving a model requires expensive hardware, power and memory, the very input in shortage and rising.
Hardware vendors are betting efficiency solves it. Nvidia says its next platform generation delivers another leap in efficiency per token processed. Even so, analysts point to the difficulty of closing the gap at the current rate of price decline.
In the same period, OpenAI cut prices across its API line citing efficiency gains, and Meta opened the weights of its latest model. Different moves with one shared effect: they pull down further the price the market will accept.
For anyone buying AI as an input, the short-term effect is good and the medium-term one deserves attention. A price falling on competition rises again when competition thins, and depending on a supplier whose margin doesn't close is a continuity risk, not a cost one.
