Among the results published in late May, a model reached 57.9% on a test considered hard, with the important caveat that the number holds for runs with tool access.
The distinction between measuring with and without tools is what usually gets lost in coverage. A model with access to search, code execution and document lookup solves problems it wouldn't solve alone, and that's good.
It's good because it better describes real use. Nobody puts a model to work without giving it access to what it needs to consult, so measuring with tools comes closer to production conditions.
The caveat is about comparison. Numbers measured with tools don't compare with numbers measured without, and the configuration used is rarely detailed in launch material.
For anyone assessing suppliers, the practical consequence is familiar: the only reliable comparison is internal, running your own operation's task with the same tools available to both candidates.
