AMD acquired Taalas, a company that bakes a model's weights directly into silicon, a technique promising to raise execution performance by an order of magnitude or more. The announcement came at market close and is framed as another attempt to challenge Nvidia's dominance in AI hardware.
The idea trades flexibility for efficiency. A general accelerator runs any model because it holds weights in memory and loads them as needed. A chip with the model etched in gives up that generality and, in exchange, removes much of the traffic between memory and processor, which is exactly where time and power are lost.
The limit of the approach is obvious and well understood: an etched model doesn't change. Switching versions means switching chips, which only makes sense for stable models with heavy, predictable use, not for anyone still deciding which model to run.
The deal rhymes with what Anthropic announced the same week in assembling a team to design its own chips. Both start from the same arithmetic: when inference spend becomes the largest line in the budget, custom hardware stops being a manufacturer's luxury.
For the ordinary buyer, the practical effect is slow to arrive. What changes now is the expectation: efficiency per watt has become the field where the competition actually happens.
