Anthropic launched Claude Opus 4.8 in the period, scoring 69.2% on a software engineering benchmark using real repository problems, above the previous version's 64.3% and ahead of the result attributed to a competitor on the same test.
The type of test matters more than the number. A benchmark built on real repository problems requires the model to understand an existing codebase, locate the relevant section and produce a fix that passes the project's own tests, which is very different from solving an isolated exercise.
Alongside it came features aimed at continuous work rather than question and answer, including a more intensive execution mode and dynamic workflows inside the coding tool.
The company also announced separate monthly credits for agent builders, available from mid-June, creating a dedicated usage pool beyond the regular plan.
That last decision says something about the consumption pattern that emerged in the period. People building agents spend differently from people chatting with an assistant, and charging both from the same bucket ended up penalising precisely the more advanced use.
