The 83.4% result on a computer task benchmark is high and deserves the caveat the test's nature imposes: measurement happens in a controlled environment, and real interfaces change.
That's where agents of this kind fail most visibly. A button moves after an update, a window opens out of order, an unexpected notice covers the expected element, and the system that learned the path keeps clicking where it shouldn't.
The aggravator is that this error is usually silent. The agent doesn't notice the action had no effect and continues the sequence, producing a result that looks complete and is wrong.
It's the same behaviour observed in a product launched weeks later, with reports of execution continuing through ambiguity without stopping to confirm.
For anyone considering an agent that operates interfaces, the design question comes before choosing a supplier: how does the system know the action worked, and what does it do when it can't confirm.
