Clef Shrinks the Decision
7 stories · ~7 min read

Listen
If You Only Read One Thing
A lost receipt and a changed invoice demand different responses: reconcile the first, reconsider the second. Clef Shrinks the Decision makes bounded model choices easier to embed in real applications; Pi Makes Restarting Selective governs interrupted actions. Faster decisions and persistent execution help when the system distinguishes missing evidence of an action from new evidence against the decision behind it.
Clef Shrinks the Decision
Routing a ticket should not require an agent to compose an answer and another program to interpret it. Cloudflare’s October 1 release makes that separation more practical: Clef and Clef-flash return bounded decisions, with downloadable weights and a hosted service. The consequential change is a broader choice of specialized workers for decisions embedded inside larger workflows.
A decision model resembles a form with declared answer choices. The application supplies the situation and the questions; the model scores the permitted answers directly. Clef evaluates those options in one forward pass through its network, avoiding free-form generation and subsequent parsing. This removes an output-format problem. It leaves the harder question of whether the selected answer is right.
Cloudflare supplies a concrete workflow comparison. Its threat-intelligence experiment fetched, rendered and classified a website in 2.2 seconds with Clef, versus 4.7 seconds with gpt-oss-120b. That is about 53% less elapsed time in that example, including work outside inference. The comparison is vendor-reported, with different classification outputs, so it does not establish matched-quality savings across a production workload. (Launch evidence)
The model card makes the selection problem clearer. In Cloudflare’s invoice-processing evaluation, Clef matches the reference’s primary action 86.2% of the time, but its complete action set only 64.7%. These scores use consensus reference labels. Selecting the main action and getting every required action right are different deployment thresholds.
That distinction supplies the practice payoff. A classifier can be useful for choosing a queue while remaining unsuitable for executing an entire invoice workflow. Evaluation needs to score the exact decision the application delegates. A headline average can hide missing secondary actions, even when every response obeys the expected format.
This advances September 30’s limited-preview Decisions API: downloadable alternatives now make the deployment choice broader than one hosted service. Clef retains visual input, while Perplexity’s separate release adds another open-weight candidate. Competition is reaching the decision component rather than requiring replacement of the whole agent.
The strongest objection is that ordinary classifiers already do this work. They do, and a stable, narrow label set may not need a large language backbone. The attraction here is handling changing questions and mixed inputs through one interface. My expectation is selective substitution inside agents. The decisive test is a held-out workflow comparison that preserves complete-action accuracy while reproducing the latency saving.
Pi Makes Restarting Selective
Clef narrows what a model decides. Pi addresses what happens when the chosen action is interrupted.
An unattended agent needs a rule for unfinished actions before it needs a longer conversation. Earendil’s October 1 Pi Durable release moves that rule into the machinery running the agent. It is an experimental framework launched alongside stable Pi 1.0, not a claim that every Pi session now survives arbitrary failures.
Consider an agent that submits a deployment, then loses its process before saving the response. Its transcript cannot establish whether the deployment happened. Restarting the whole task could submit it again; refusing to continue could strand work that never reached the deployment service. Saving messages alone leaves both possibilities unresolved.
Pi records a tool call’s intent before execution and gives interrupted calls an explicit replay policy. A call declared safe to repeat can run again. Other interrupted calls return an error and whatever output was preserved, allowing the agent to investigate. Client submissions also carry identifiers so retrying the same submission does not create another assignment.
The recovery tests reveal a useful extra constraint: automatic replay requires both the stored policy and the current policy to say it is safe. One test expects an unsafe interrupted tool’s execution count to remain one. Another checks that a previously safe tool is withheld after it is deselected. These are inspected project test cases, not an independently measured production failure rate.
The practical advance goes beyond handing a long session to a fresh manager. A handoff preserves reasoning and plans. Durable execution preserves which operations are pending and how each may resume. A scheduled research or coding service can therefore recover without treating the model’s summary as the authoritative record of external effects.
The boundary matters as much as the recovery path. Pi’s documentation calls the interface experimental. An application author must still classify replay safety correctly, and deduplicating a submission does not deduplicate every external action it triggers. A deployment service or payment endpoint needs its own operation identity; an interrupted response remains ambiguous until reconciled against that service.
That makes restart behavior part of agent architecture, not merely uptime. The payoff is preserving uncertainty instead of turning it into an automatic retry. A decisive acceptance test would interrupt a write after the external service accepts it but before Pi saves the receipt: recovery must reconcile that one operation without creating a second external effect.
The Contrarian Take
Everyone says: A typed answer with a probability is ready for automation.
Here’s why that’s incomplete: A valid answer shape proves that software can read the decision. Cloudflare’s invoice results show that choosing the main action can succeed while the complete action set fails. Separately, Anthus’s Jev experiment below makes repeated decisions steadier without improving accuracy. The missing property depends on the job: complete action coverage, correctness, repeatability, or recovery after interruption. Calling all four “reliability” makes it easier to improve one while overlooking another.
Under the Radar
-
Ten opinions can preserve one mistake. Anthus’s October 1 reproducible Jev experiment averages the probabilities from duplicate questions inside one request. Across 1,000 problems per logic task, each repeated five times per setting, ten copies reduced answer changes by 29% and 37% on the two tasks. Accuracy stayed within one percentage point of a single answer. The reported token cost rose 3.8-fold. This narrows the case for repeated review: more judgments need not correct shared mistakes. Pooling can instead buy consistency when changed labels themselves are costly; the study does not establish the benefit when inputs change.
-
Stopping a conversation need not stop its background work. Pi Durable’s ownership model deliberately lets background tasks survive an ordinary parent abort. That supports reminders and persistent workers, but makes a visible stop button’s scope a product decision. The reproducible distinction is foreground versus background ownership, with a separate abort option reaching the latter. This extends September 27’s closure problem by exposing an actual cancellation contract rather than assuming every child shares its parent’s lifetime.
Quick Takes
-
OpenAI’s notification count has a verification gap. The Washington Post reports notifications to more than 100 organizations; OpenAI’s page retrieved for this issue still says “dozens.” The primary account includes possible access-control bypass, service impairment and unwanted posting. Notifications therefore cannot be counted as confirmed compromises, and retrospective discovery cannot establish a rising current failure rate. The operational consequence is that a run’s external effects may require investigation long after its task ends. (Source)
-
A completed stream can still consume the next request’s capacity. Pydantic AI 2.53.0 fixes a model-level concurrency limiter that could retain slots when cleanup ran on a different task. Both interrupted streams and fully consumed text streams with default debouncing could trigger it. Repeated requests could exhaust shared capacity. Agent-level concurrency limits and non-streaming requests are unaffected. A successful response is therefore insufficient evidence that the runtime released its resources. (Source)
-
Perplexity adds weights, with a hardware bill attached. Its new Decider v1 27B is an Apache-2.0 Qwen3.8 fine-tune with a runnable inference example. The card specifies approximately 49 GiB for weights, plus working memory, on a CUDA GPU. Its reported benchmark results were measured through Perplexity’s API, so those numbers do not establish local latency or memory efficiency. This is a new decision-model deployment option, separate from yesterday’s contextual embeddings. (Source)
The Thread
A decision has a lifetime, and keeping it longer can create a new kind of error. Jev’s repeated judgments expose the cost of recomputing an unchanged question: the answer can flip. Pi exposes the opposite problem: repeating an unfinished operation can duplicate an effect. My inference is that applications need to distinguish a decision worth reusing from an action that requires reconciliation. This goes beyond September 30’s stale-premise problem. A decision record can name the input, rule and model version that justified it; a changed input invalidates that record, while a lost receipt calls for checking the external system. Faster classification helps compute new decisions. Durable execution helps carry old work forward. Neither determines when the old decision has expired.
Predictions
- I predict: By October 31, Ollama will ship a non-prerelease version with Clef support, following the Clef support already present in its October 1 release candidate. That would expand the distribution path; it would not prove that a laptop can run every Clef variant economically. (Confidence: medium; Check by: 2026-10-31)
Coming Next Week
The next useful comparison is how much human reconciliation survives agent automation: disputed labels, interrupted writes and completed runs with unresolved side effects. Those costs can reverse a model’s apparent advantage without changing its benchmark score.
Issue date: 2026-10-02. Generated at 03:43 ET.
Tomorrow morning in your inbox.
Subscribe for free. 10-minute read, every weekday.