The announcement of a model in private beta restricted to companies in the announcer's own group raises a question of method that matters beyond that case.
Testing in real operations has a clear advantage over bench evaluation: problems surface in use, with real data and real consequences, rather than in a scenario built to measure.
The limit is the sample. Companies in one group share culture, infrastructure, tooling and problem types, producing a favourable and unrepresentative picture of what happens when the model meets the rest of the world.
Any test run by someone with an interest in the result carries that limitation, and it applies to evaluation numbers published at launch: the benchmark is chosen by the publisher and comparisons rarely use identical configurations.
For anyone assessing a supplier, the practical answer is familiar and laborious: measure on your own workload, with tasks your operation actually runs, because the ranking usually changes once testing leaves the marketing material.
