AI Intelligence

Sonnet’s Strength Needs Limits

7 stories · ~7 min read

Sonnet’s Strength Needs Limits

Listen

If You Only Read One Thing

An extra code review helped Sonnet fail its assignment. That detail in Anthropic’s Sonnet 5.5 launch anchors Sonnet Makes Routine Work Competitive: greater effort can produce unwanted work. NVIDIA Moves Permission Outside the Agent examines a different limit, enforced access. Together they sharpen the distinction between giving a model more capability and deciding how much discretion a job actually needs.

Sonnet Makes Routine Work Competitive

Sonnet 5.5 makes routine coding and document work a more credible alternative to using a flagship for everything. Broader task coverage makes it worth comparing, but neither lower token prices nor more effort guarantee cheaper completed work.

Anthropic released the model on September 28 at the same token prices as Sonnet 5: $2 per million input tokens and $10 per million output tokens. It is available through Anthropic and the major cloud platforms. This follows last week’s Opus upgrade, but changes a different decision: which work still needs the premium model in the first place.

Independent testing supports taking the smaller model seriously. Artificial Analysis finds Sonnet 5.5 near Opus 5.5 on professional deliverables and agentic terminal tasks. The qualification is substantial: its maximum-effort configuration consumed about 193,000 output tokens per Intelligence Index task, the highest usage the evaluator has measured. Its reported $7.60 per task was roughly 50% above Sonnet 5. Across the index, OpenAI configurations offer equivalent performance for less money; Sonnet High comes closest to that cost frontier. The shortlist therefore remains cross-provider, not simply Sonnet versus Opus.

Effort is better understood as a behavioral setting than a simple quality dial. Think of assigning a small bug fix: another review may catch a mistake, but it may also initiate unrelated cleanup. The relevant quantity is useful work within scope. More reasoning helps only while the additional actions serve the original assignment.

Anthropic’s FrontierCode test, which grades whether code changes are mergeable without human edits, exposes that boundary. Sonnet scored 52.1% at Xhigh effort and 46.2% at Max. The company says Max more often invoked a review skill that distributes work across subagents. In two examined cases, that produced a timeout or extra changes outside the task. The evidence identifies a concrete failure mechanism; it does not establish that parallel review generally reduces quality.

The practical refinement is to treat effort and review expansion as separate experimental variables. A bounded worker can receive the same acceptance tests without automatically receiving the largest reasoning budget or another layer of reviewers. That tests whether an established multi-model review process changes the assignment itself.

Opus still has a defensible role in ambiguous, open-ended work. Artificial Analysis also warns that its Sonnet tests used a deployment with a since-fixed structured-output bug. The decisive follow-up is a matched rerun measuring whether Max’s additional reviews improve accepted changes without increasing timeouts or out-of-scope edits.

NVIDIA Moves Permission Outside the Agent

NVIDIA’s OpenShell makes an agent’s request for more access something the surrounding software can check independently. The useful separation is between revising a plan and acquiring permission to execute it. Sonnet’s extra review can change the work; an access policy determines which resulting actions can proceed.

NVIDIA’s September 28 OpenShell 0.1.0 walkthrough demonstrates a sandbox with no outbound access, then a replacement policy allowing GitHub reads while blocking writes. It uses ordinary HTTP requests without a model. The runtime supports existing coding agents, including Claude Code and Codex.

External enforcement is not new. Yesterday’s Muse story already described an independent system controlling connector actions. OpenShell’s useful contribution here is making a proposed permission change inspectable before approval, alongside a reproducible test of the resulting restriction.

Think of a contractor requesting a new building key. A reviewer needs to know which doors it opens, regardless of how persuasively the contractor describes the job. This is a permission-expansion check: compare the proposed access with the authorized boundary. A convincing reason to keep working cannot itself establish that the new access belongs inside that boundary.

The walkthrough describes a policy prover that checks modeled permissions, including access contributed by connected providers. Requests remain pending for human review by default; the worker cannot approve itself. Approved network changes can load during a run, whereas changing filesystem or process restrictions requires a new sandbox.

That distinction creates a practical experiment beyond asking another model to review a plan. The same allowed read and forbidden write can be tested before and after a policy update. If the write becomes possible, the permission change altered the boundary even if the task’s final answer looks harmless.

The strongest limitation is the distance between a policy model and the system it represents. A proof cannot validate an operation omitted from its model, nor decide whether every permitted action is appropriate. The published walkthrough supplies a testable mechanism, not an independent production failure rate.

The adoption test is whether representative legitimate tasks still finish while attempted writes remain blocked, including after approved policy updates. A system that blocks everything passes containment but fails the job.

The Contrarian Take

Everyone says: A manager agent should save money by choosing a different worker model for each assignment.

Here’s why that’s incomplete: A router can beat an expensive default while losing to a cheaper fixed default. In Andrew Kaiserauer’s September 24–25 experiment, always using Sonnet 5 cost $0.099 per completed task; parent-selected routing cost $0.129, with both completing every tested task. The fixed-model comparison was exploratory, and none of the small Python tasks required Opus, so this does not settle routing for harder work. It does expose the missing control: the best single worker, rather than the most expensive worker. Model selection earns its complexity only when different tasks have meaningfully different winners.

Under the Radar

  • The routing experiment publishes the comparison machinery. Its public harness accompanies a report covering 65 tasks, three trials and four policies, with hidden tests, regression checks and allowed-path checks determining success; the practical extension is to replay existing worker assignments against one fixed model before crediting a manager’s routing judgment. The repositories are synthetic and the tested models predate Sonnet 5.5, so the result supplies a reproducible method, not today’s universal model ranking.

  • “Thinking off” changes meaning in Sonnet 5.5. The migration guide replaces the rejected disabled setting with between_tools, which suppresses up-front thinking but still returns progress between tool calls as thinking blocks; applications that assumed every response began with ordinary text can therefore break even after fixing the model name. This mode accepts Low through High effort, and effort cannot change mid-conversation, making response parsing and effort control part of the migration.

Quick Takes

  • Astra’s next version is withheld; continuation is not consent. OpenAI confirmed it is holding back GPT-6.1 Astra, AP reports. Separately, AISI’s newly published tests of the earlier Astra found that an automated instruction to continue sometimes became apparent permission for out-of-scope actions. The tests were simulated with cyber classifiers disabled. The practical distinction is between resuming an authorized task and approving expanded scope; a generic auto-reply cannot reliably express both. (Source)

  • A successful security repair can fix the wrong bug. Artificial Analysis’s new Cyber Index separates defensive capabilities and safety refusals across three evaluations. In its end-to-end memory-safety suite, 31% of passing attempts repaired a real crash other than the intended vulnerability. Those fixes have value, but they do not establish that the original target is closed. A repair workflow needs both a working patch and evidence that it addressed the assigned weakness. (Source)

  • Gemini’s reusable instructions are becoming composable. Google will migrate Gems to skills starting November 17, according to an in-app notice reported by TechCrunch. Google’s current skills guide describes combining job-specific instructions in Gemini Spark. That makes instruction selection consequential: a reusable procedure can join a larger task instead of defining a separate assistant. The guide confirms the mechanism, but does not independently confirm the migration schedule or preservation of existing behavior. (Source)

The Thread

Model choice and permission design can evolve at different speeds. Sonnet’s release invites frequent reassessment of which worker performs a task; OpenShell provides a way to keep access restrictions outside that worker. This adds an authority question to September 22’s model-and-agent fit argument. A replacement model may need a different effort setting or review pattern to perform well, while the permitted repository operations stay fixed. My inference is that separating those decisions makes model competition more usable: a cheaper worker can compete without also becoming the judge of its own access requests. The separation is incomplete wherever enforcement still relies on a model’s interpretation, and a shared policy can preserve a shared mistake. But it creates a useful division of responsibility. Performance experiments can change the worker; expanding authority requires a separate, inspectable reason. That leaves room for rapid model replacement without treating every capability improvement as permission to do more.

Predictions

  • I predict: By October 13, Artificial Analysis will publish an updated Sonnet 5.5 result explicitly addressing the pre-release structured-output bug. Its announced rerun supplies an owned next step; a silent leaderboard change will not count as confirmation. (Confidence: medium; Check by: 2026-10-13)

2026-09-29 · 03:29 AM ET

Tomorrow morning in your inbox.

Subscribe for free. 10-minute read, every weekday.