Sonnet 5.5: Higher Effort Isn’t Always Better

Sonnet 5.5 has a useful warning tucked into its launch announcement: on one coding benchmark, turning effort up to Max produced a lower score than Xhigh.
Anthropic’s September 28 release explains why. FrontierCode asks whether a change can be merged without human edits. Helpful changes outside the request still count against it. At Max, Sonnet more often invoked a review skill that divides work among subagents. In two cases Cognition examined, that extra review led to a timeout or edits beyond the task. These are vendor-reported observations from a particular evaluation, rather than a rule that Max is worse. Anthropic’s announcement, footnote 2
The useful question is what to do with that information. When an agent’s output disappoints, “let it think harder” is only one possible response. Sometimes the missing instruction is where the work should end.
The button works. The patch still needs another decision.
Consider a small, fictional checkout. The button says Pay securly. The request is straightforward: correct it to Pay securely, preserve the payment behavior, and leave the settings alone.
For this article, we made two patches and ran them locally. These are deliberately authored examples; neither was produced by Sonnet. They illustrate a completion check, without pretending to compare effort settings.
Patch A changes the label. Patch B changes the same label and also gives the page a purple background and a friendlier receipt setting. A designer might prefer the new color. The receipt change might be worth discussing. Both are additional decisions.
The functional check accepts both patches: the label is correct, and one click still invokes one payment request for US$48.00. The task check accepts only A, because B also changes the settings file. The original, misspelled version fails the label check.

Original fictional replay, run October 5. No model comparison or real payment. The fixture, completion contract and results are retained with the production source.
Nothing in this example requires us to call B bad code. The problem is that someone must now decide whether the new settings belong in this release. A task that was ready to finish has acquired another review.
That is easy to miss when the completion message lists only successes: fixed the label, preserved behavior, improved the page. The changed-file list tells a different story. It asks the owner to approve work they never requested.
More review can create more work
A review step can uncover a real defect. It can also produce optional improvements: a clearer name, a better layout, another test, a refactor nearby. Each suggestion may be reasonable on its own.
The boundary gets crossed when “consider this” silently becomes “implement this.” The agent’s expanded result can pass the original test while the whole patch becomes harder to accept. Running the same functional check again won’t reveal that distinction; the result already passed it.
This is the mechanism the Sonnet footnote makes worth examining. More available effort creates room for further investigation and review. The task still needs a separate definition of which findings require action. A code review can return a recommendation without automatically widening the patch.

Original illustration. Keep optional improvements as recommendations until they belong in the task.
For the checkout example, a usable request is short:
Fix the button label from “Pay securly” to “Pay securely.” Preserve the amount, currency and click behavior. Leave the existing settings unchanged. Check the label and payment behavior, inspect the diff, and return the patch. List unrelated suggestions separately.
That instruction gives the agent something more useful than a vague request to be thorough. It defines the requested result, what must stay stable, and the evidence needed to finish.
Diagnose the failure before moving the slider
If the agent stays within the requested task but reaches a wrong answer, additional reasoning may help. A difficult dependency, competing explanations or an unfamiliar algorithm can justify more investigation. Check whether the next attempt actually resolves the uncertainty.
If the answer is correct but the patch includes unrelated work, clarify the boundary. More thinking alone doesn’t settle whether those changes belong. Separate necessary fixes from optional recommendations and review the diff against the request.
If a clearly bounded task remains unsolved after relevant checks, reconsider the model, tools or available context. Repeating the same attempt at a higher setting can be less useful than supplying the missing file or an accurate example of the expected behavior.
These are different failure types. Treating every one as “insufficient effort” makes the setting carry decisions it cannot make by itself.
Some tasks need room to discover their scope
The checkout is deliberately simple. “Find why checkout fails intermittently” is a different assignment. The cause may be in payment handling, retry logic or configuration. A rigid one-file limit could prevent the agent from finding the problem.
For that kind of work, authorize investigation broadly and define how it becomes a change: establish a reproduction, explain the cause, identify the necessary edits, and flag any consequential expansion. Give high effort a problem worth investigating. Keep the completion rule appropriate to that problem.
Even the small typo task has an exception. If the requested edit exposes a dependency that must change for the fix to work, the agent should explain it. A stopping rule should prevent casual expansion while allowing a necessary finding to reach the owner.
Choose the finish line with the effort level
Sonnet 5.5’s launch gives builders another reason to pay attention to effort. Its footnote also gives us a concrete reminder that the score belongs to a task with an acceptance standard. More activity is useful when it moves the result toward that standard.
For routine repairs, make completion visible: the requested change works, the protected behavior holds, and the diff contains the work that was agreed. For an investigation, make the unresolved questions and permitted expansion visible instead.
The next time an agent returns an impressive patch that still cannot be merged, inspect what made it unacceptable before asking for a stronger attempt. It may need more reasoning. It may need better context. Or it may already know how to solve the problem and need a clearer instruction to stop.
