Sol Challenges Premium Reasoning
7 stories · ~7 min read

Listen
If You Only Read One Thing
A read-only agent can still change what happens next. Its notes survive to inform future authorized work, a boundary explicit in Dots’ permission rules. That is the tension between Sol Challenges the Expensive Default and Dots Separates Discovery From Action: cheaper workers make execution easier, while persistent assistants can keep supplying new jobs from old premises that nobody has corrected.
Sol Challenges the Expensive Default
GPT-6.1 Sol makes a credible case for moving ordinary work off the premium model. The consequential change is how little measured capability that move now appears to sacrifice.
OpenAI’s September 29 announcement places Sol’s standard input and output token prices at one-fifth of Astra’s and makes it available across paid ChatGPT plans and the API. That is a rate comparison, not a completed-task saving: the amount of reasoning, failed attempts and subsequent repairs still determine the bill.
Independent evidence supports the broader direction. Artificial Analysis’s Sol evaluation gives maximum effort a rounded intelligence score of 52, against Astra’s 53. Its average evaluation-task cost falls from $3.26 to $0.72, about 78%. The index combines different kinds of work, so the one-point gap is not a one-percentage-point coding accuracy gap.
Think of this as paying for an exception desk. Most requests go through a cheaper service; difficult cases justify specialist attention. The mechanism works only when the ordinary service handles the actual workload and failures can be recognized. Otherwise, escalation adds a second bill after the first attempt. The useful distinction is default versus exception, with evidence deciding where a task belongs.
A small, task-specific evaluation offers a narrower claim. The Explicit Edit dataset records Sol 6.1 at low effort completing all 226 tasks exactly on the first attempt in Pi’s default harness. These are prescribed file changes, including Unicode and large-file cases. They test execution of an already-specified edit, not discovery of the right software design.
That adds something to the existing practice of delegating work and reviewing it with another model: a cheap worker can be assigned a demonstrably narrow contract before the expensive reviewer enters. The strongest objection is that real changes seldom arrive fully specified. A perfect edit score says little about understanding an ambiguous requirement, and the broader index can conceal a critical weakness in one specialty.
Yesterday’s Sonnet analysis examined extra effort expanding the job. Today’s result extends September 23’s Sol cost comparison to a new version and a prescribed-edit workload. The adoption test is whether Sol preserves first-pass acceptance on a fixed set of previously expensive tasks after repair and escalation costs are included.
Dots Separates Discovery From Action
An always-on assistant need not have always-on permission to change things. That distinction is the most consequential technical detail in OpenAI’s Dots launch: noticing possible work and executing assigned work run under different restrictions.
Dots launched September 29 for Pro and Business Premium in eligible markets, with an administrator-enabled beta for Enterprise, Edu and Healthcare. The rollout makes persistent assistance a product surface; it does not establish dependable performance on every unattended task.
The important mechanism is the discovery-to-action boundary. Imagine an assistant reading customer feedback and noticing repeated complaints. Finding the pattern does not itself authorize contacting customers. OpenAI’s safety FAQ says proactive research can read permitted sources and keep private notes, but cannot directly message people, edit plugin content, or control a browser or computer.
Assigned work follows separate action rules, including when delegated or continued in the background. Custom Rules cannot disable the independent review system or relax proactive-research restrictions. “Background” describes when work runs, not one uniform permission level.
The practical payoff extends familiar scheduled agents. A schedule says when to inspect; this design also distinguishes an observation from a mandate to act. A monitoring agent could prepare evidence about an overdue invoice without acquiring standing permission to send it. That is a workflow implication, not a measured productivity gain.
The strongest counterargument is that discovering work already changes the system. The FAQ says individual dot memories cannot currently be viewed, deleted or directly edited. Disconnecting a plugin stops new access but leaves previously retained context. Correcting a conversation therefore need not mean removing the premise that informed it. The documentation does not establish how reliably later corrections override earlier context.
Monday’s Muse case separated permission from factual accuracy; yesterday’s OpenShell story tested access expansion. Dots adds persistence: an old inference can inform a later authorized job. A useful test is whether an assistant still proposes the same mistaken invoice after the underlying record is corrected, without deleting the dot’s entire context.
The Contrarian Take
Everyone says: Once near-frontier intelligence is cheap enough, a premium model becomes unnecessary for ordinary agent work.
Here’s why that’s incomplete: Sol’s $0.72 average evaluation-task cost is meaningful evidence for substitution, but an average does not identify the exceptions. The exact-edit result is stronger evidence for one bounded operation because the desired output is known in advance. An assistant deciding which neglected task deserves attention has no equivalent byte-for-byte answer. My judgment is that falling model prices favor more specialization inside a workflow: economical execution where success is explicit, stronger reasoning where interpretation remains difficult. As yesterday’s routing comparison showed, an expensive coordinator could consume the saving before the cheaper worker starts; the coordinator needs its own measured contribution.
Under the Radar
- A perfect score can expose the harness instead of the model. Two low-effort Sol configurations in the Explicit Edit results both finish all 226 tasks exactly, but report about 1.97 million versus 498,000 total tokens, roughly a fourfold difference. These are different harness configurations, not a controlled model comparison; the dataset’s zero cost field does not mean free inference. Once correctness saturates, context and tool overhead become the remaining comparison, which extends model selection into the surrounding execution loop.
- Shared work now has a persistent editing surface. ChatGPT Space and Pages give teams and agents shared knowledge and editable documents, with Space available on Pro, Business and Enterprise. The architectural opportunity is to keep a project’s current state outside any one conversation. That also changes how mistakes spread: an incorrect shared document can inform several later tasks. Collaboration features establish access and editing, not proof that every agent used the latest revision or interpreted it correctly.
Quick Takes
- GLM crosses a consequential cyber threshold. Anthropic’s September 29 tests found GLM-5.3 completed 50 of 410 exploit attempts on a browser-engine benchmark, versus Mythos Preview’s 56. Separate tests found complete control-flow hijacks in 4% of 100 sampled tasks. These are sandboxed evaluations by a competitor, not observed attacks. The deployment implication is still substantial: defensive testing can no longer assume advanced exploit capability resides only behind restricted flagship APIs. (Source)
- Ultrafast buys generation speed, not an eightfold shorter project. Astra’s premium tier is now available, with OpenAI advertising up to eightfold faster generation in Codex and sixfold in the API; Sol’s version is still forthcoming. That helps when waiting for model output dominates. Tests, tool execution and human review remain on their own clocks. The launch-stage demonstration establishes a faster interaction loop, not equivalent end-to-end throughput across production workloads. (Source)
- Decisions gives a small model a bounded answer space. OpenAI’s limited-preview API uses Luna to answer predefined questions with finite choices from text or image context. Classification and routing can therefore avoid generating an open-ended response. This extends the typed-decision pattern covered with Jev on September 17, but no matched accuracy, calibration or price comparison was established in the launch material. A bounded answer can still be the wrong answer. (Source)
The Thread
Cheaper models make the location of difficult judgment more visible. Sol’s reported results support testing specified work on a less expensive execution path. They do not show that choosing useful work has become equally cheap or reliable. A workflow contains several distinct jobs even when one agent presents the result.
Dots separates background discovery from action, extending yesterday’s worker-versus-permission distinction. But its retained context introduces another dependency: the premise behind a future assignment can outlive the source connection. A permitted action can therefore inherit an obsolete observation without crossing an access boundary.
September 26’s assignment-quality argument concerned translating human preferences into a target. Today’s additional problem is revising that target as evidence changes. Space can spread a corrected project record, but shared editing alone does not prove that every agent has abandoned the older premise.
My inference is that persistent assistants need a measurable correction path. The relevant outcome is whether a corrected fact stops generating the same proposed work across later sessions. Faster, cheaper execution increases the cost of failing that test: an obsolete assignment can keep returning even when every attempt is inexpensive and properly authorized.
Predictions
- I predict: By October 14, OpenAI will make GPT-6.1 Sol Ultrafast available on at least one production surface, beyond a launch announcement. The company has already shipped Astra’s tier and explicitly committed to Sol next; that is firmer evidence than an inferred ecosystem roadmap. (Confidence: medium; Check by: 2026-10-14)
2026-09-30 · 03:33 AM ET
Tomorrow morning in your inbox.
Subscribe for free. 10-minute read, every weekday.