Capacity shortages at the leading cloud providers began reorganising large companies' AI strategies, accelerating diversification into custom chips, alternative clouds and efficiency measures, with a direct effect on smaller companies' access to frontier models and on deployment timelines.
The point that usually escapes is that scarcity doesn't hit everyone equally. Those with consumption history and long contracts keep allocation; those starting out join a queue where the criterion is volume already contracted.
That changes what a new company can promise. Commitments on response time and volume depend on contracted capacity, and contracted capacity depends on availability nobody guarantees in writing.
The responses that appeared in the period were the same at every scale of operation, varying only in budget: more than one supplier, cost-based routing between models, and open models on your own hardware for the predictable part of the load.
It's the software version of the lesson hardware learned with components: a single source isn't a choice, it's a risk taken without compensation.
