trampolim.net
PT EN

Trampolim · Technology weekly

The Week in Tech

Issue 05Week of May 25–31, 202616 stories
AI

Computer-use benchmarks join coding in the scorecard

The 83.4% figure comes from a test requiring operation of a graphical interface, not just writing code.

Computer-use benchmarks join coding in the scorecard
AI · May 25–31, 2026

Among the results published at the late-May model launch, one stands out for measuring a different skill: 83.4% on a benchmark of tasks executed in a computer environment with a graphical interface.

Operating a graphical interface is a distinct problem from writing code. It requires interpreting what's on screen, deciding where to click, handling windows that open out of order and recognising when an action didn't have the expected effect.

That's precisely where agents fail most visibly, and the reason is structural: interfaces change, buttons move, and a system that memorised the path errs without noticing it erred.

Progress on that benchmark explains the commercial interest in agents that operate applications, which appeared in a product launched weeks later with a monthly price and a promise of autonomous work.

For anyone considering that kind of agent, the practical question stays the same: what happens when it errs, who notices, and how much it costs to redo what was done wrong before anyone spotted it.

Book a call