Cheap Judges Need Calibration
7 stories · ~7 min read

Listen
If You Only Read One Thing
A judge can become more trustworthy without catching more mistakes. Microsoft-Decision-1 makes automated judgment cheaper; Typed Evals’ calibration experiment shows why confidence needs its own test. The distinction matters when a score decides whether an agent continues, escalates, or stops. Cheaper decisions expand automation only if the number attached to each decision means what the surrounding software assumes it means.
Microsoft Makes Judgment Cheap
Microsoft-Decision-1 makes it economical to put a model between ordinary workflow steps. The consequential question is whether those inexpensive judgments are good enough to control what happens next.
Microsoft’s October 9 release is a decision-scoring model derived from Qwen3.5-9B. It costs $0.042 per million input tokens, with no output charge. Microsoft reports leading accuracy across 36 benchmarks containing nearly 150,000 questions and a 35-fold speed advantage over GPT-6 Sol in its tests. Those are vendor results, not a general speed guarantee.
A decision model resembles a dispatcher with a fixed menu. The application supplies the situation and possible destinations; the model assigns probabilities and selects an option instead of composing an explanation. A support request can become “billing,” “technical,” or “account.” Removing open-ended generation makes these repeated choices cheaper, but the dispatcher cannot select a destination the application forgot to include.
Independent evidence adds a different comparison. Benchmark Heaven’s October 10 Foundry evaluation places Decision-1 sixth on its hosted composite ranking, which combines quality, calibration, speed and cost. Its displayed median endpoint latency is about 460 milliseconds. That is a hosted-service measurement, not the time for the model’s computation alone. The two evaluations use different workloads, so this is not a matched rerun of Microsoft’s test. It shows a useful competitor without independently establishing the headline speed advantage.
Serving conditions matter particularly when the model itself is fast. An agent making sequential decisions pays for the complete round trip each time. Microsoft illustrates the accumulation with twenty decisions adding 100 milliseconds apiece: two seconds of extra waiting. A very cheap classifier can still make an interactive workflow feel slow; low token prices do not remove sequential delays.
The strongest case for adoption is bounded substitution: scoring familiar choices inside an existing application. This extends October 2’s decision-model argument with a new hosted competitor and independent endpoint evidence. It does not justify replacing the reasoning that establishes which choices are legitimate. The application still owns escalation and execution permissions.
The deciding comparison is a fixed workload at a fixed tolerated error rate, including requests the endpoint cannot accept. If Decision-1 lowers total decision cost without increasing unresolved or wrongly routed cases, the cheap-dispatcher thesis holds; a lower token bill alone cannot establish it.
Confidence Needs Its Own Test
Calibrating an AI judge can improve the meaning of its scores while barely changing which answers it gets right. That makes calibration useful for allocating review, but a poor substitute for improving the judge.
An October 10 practitioner announcement brings a concrete example: Typed Evals, a Python evaluation toolkit. Its published experiment is dated September 27, not this weekend. The timing matters because this is newly surfaced practice evidence alongside Microsoft’s launch, rather than a new breakthrough claim.
Calibration is the same test applied to a weather forecast. Among days assigned an 80% rain probability, roughly eight in ten should be rainy. For an AI judge, the corresponding event is human acceptance of an answer. A score becomes useful for deciding how much uncertainty to tolerate only after it has been compared with those observed outcomes. Call this confidence with a denominator.
The TRIVIA+ experiment evaluated 645 answers against their source articles. It reused the same saved Jev judgments before and after calibration. The average gap between grouped confidence scores and observed outcomes fell from 0.0982 to 0.0313. Classification accuracy, after selecting cutoffs on separate validation data, stayed around 72%.
The unchanged accuracy is the useful result. The adjustment supplied no new evidence about the articles and required no additional judge calls. It changed how the existing scores should be interpreted. Leaving the default cutoff unchanged actually missed more hallucinations, so calibration and choosing an acceptance threshold are separate jobs.
The toolkit’s procedure fits an adjustment from raw scores to human pass/fail labels, checks held-out examples, and saves the adjustment for later runs. In the benchmark, no identical article crossed the training, validation and test groups, reducing leakage between them. The included demonstration labels are synthetic; they exercise the machinery and cannot establish how real users judge real answers.
This extends October 5’s grader-design discussion: defining the right criterion and adding an independent reviewer still leave the numerical score unvalidated. A calibrated reviewer can help allocate scarce human attention, provided the judge, criterion and incoming work remain comparable. It cannot repair missing evidence or guarantee transfer to code review.
The adoption test is held-out agreement at the intended review threshold: does the observed failure rate match the promised confidence while preserving acceptable work coverage? Better-looking probabilities without that check are merely different numbers.
The Contrarian Take
Everyone says: A decision model that gives stable answers is a safer controller for an agent.
Here’s why that’s incomplete: Microsoft reports that Decision-1 changed its answer on only 1.3% of tested input perturbations, including reordered choices and harmless formatting changes. That measures resistance to irrelevant changes. A controller also needs the opposite property: sensitivity to relevant changes, such as a refund becoming unauthorized or a record referring to a different account. A system that consistently chooses the same wrong action can look excellent on stability alone. My proposed companion test is a paired request whose decisive fact changes; the decision should change too. This is a test design, not a reported Microsoft result.
Under the Radar
-
A handoff can disagree with the repository. The workflow project shared on October 10 runs a ground-truth script at both session start and handoff, comparing notes with git, pull requests, continuous-integration results and background jobs. The useful addition beyond a persistent task queue is explicit reconciliation before the next task is proposed. Its README supplies a reproducible procedure, not measured productivity gains; its illustrative failed-job story is not a field result. The repository makes the distinction inspectable.
-
An unavailable judgment needs its own state. Typed Evals distinguishes unavailable evaluations from ordinary failing scores: a metric’s pass value can be absent rather than false. This matters when application code treats an empty or skipped evaluation as permission to continue. A threshold comparison needs a defined result before it can justify accepting an answer. The toolkit’s contracts expose that integration boundary; they do not prove that a downstream application enforces it correctly.
Quick Takes
-
H2O supplies a local counterexample to hosted-only decisions. Its October 10 latency recheck of H2O-Lightning-4B v1.2.3 reports 205 correct answers from 231 public requests on each of two GPUs. The Apache-2.0 model runs through vLLM and a small interface shim. Startup and warmup are excluded: this is useful evidence for an already-running local service, not a cold-start or cloud-cost comparison. (Source)
-
More thinking still needs a workload-specific return. ToneBench’s October 9 Gemini 3.8 Flash comparison uses the same ten reference scripts at medium and high effort. High costs about 2.7 times as much and takes about 2.7 minutes instead of one; its writing score improves only slightly. The small evaluation cannot settle general writing quality, but it makes the additional waiting cost concrete. (Source)
-
Release watching can be a bounded agent assignment. Willison’s October 10 note records asking ChatGPT to check the Python Actions repository hourly until stable Python 3.15 appeared, then notify him. The release subsequently landed; the note does not establish that the agent reliably performed every check. The transferable detail is an observable stopping condition tied to a repository artifact, rather than an indefinite request to keep up with news. (Source)
The Thread
Cheap judgment changes the amount of work a human review queue receives. If software starts scoring every intermediate action, the operating constraint becomes how many uncertain or wrongly escalated cases people can inspect. My inference is that the important downstream price may be reviewer attention rather than inference.
Microsoft lowers the cost of producing a decision. Typed Evals shows how the attached confidence can become more interpretable without substantially improving error detection. Together, those facts suggest that a cheaper judging layer can increase review demand even while its own bill falls. Adding checks at more steps increases the number of opportunities to flag something.
This extends October 9’s visible-work argument into queue capacity. An escalation threshold has to reflect both the cost of a missed error and the capacity to investigate a flagged one. Otherwise automation can manufacture more apparently urgent work than its human fallback can finish, turning a low-priced decision into an expensive interruption.
Predictions
- I predict: By November 11, Benchmark Heaven will list at least two additional hosted decision offerings beyond the 27 on its October 10 board. Microsoft’s entry and the board’s active additions support the direction, but the rate of new submissions is uncertain. (Confidence: medium; Check by: 2026-11-11)
2026-10-11 · 03:34 ET
Tomorrow morning in your inbox.
Subscribe for free. 10-minute read, every weekday.