Elon Musk announced Grok 4.5 on June 28, built on a 1.5-trillion-parameter foundation model, in private beta across companies in his own group. Early published evaluations indicated performance near or above top competing models.
Testing in-house first, in companies with real operations, has a clear advantage over bench evaluation: problems surface in use, with real data and real consequences.
The same choice carries a limit worth recording. Companies in one group share culture, infrastructure and problem types, producing a favourable and unrepresentative picture of what happens when the model meets the rest of the world.
Evaluation numbers published by whoever ships the model deserve the usual reservation: the benchmark is chosen by the publisher, and comparisons with competitors rarely use identical configurations.
The contextual data point is scale. A 1.5-trillion-parameter base model requires training infrastructure few organisations worldwide can assemble, reinforcing the concentration running through the sector.
