trampolim.net
PT EN

Trampolim · Technology weekly

The Week in Tech

Issue 05Week of May 25–31, 202616 stories
AI

A model reaches 57.9% on a test considered hard, using tools

The detail matters: the number holds with tool access, not in a closed exam.

A model reaches 57.9% on a test considered hard, using tools
AI · May 25–31, 2026

Among the results published in late May, a model reached 57.9% on a test considered hard, with the important caveat that the number holds for runs with tool access.

The distinction between measuring with and without tools is what usually gets lost in coverage. A model with access to search, code execution and document lookup solves problems it wouldn't solve alone, and that's good.

It's good because it better describes real use. Nobody puts a model to work without giving it access to what it needs to consult, so measuring with tools comes closer to production conditions.

The caveat is about comparison. Numbers measured with tools don't compare with numbers measured without, and the configuration used is rarely detailed in launch material.

For anyone assessing suppliers, the practical consequence is familiar: the only reliable comparison is internal, running your own operation's task with the same tools available to both candidates.

Book a call