AI Intelligence

Agents Can Optimize Mistakes

7 stories · ~7 min read

Agents Can Optimize Mistakes

Listen

If You Only Read One Thing

The Grader Chooses the Winner before an optimizer takes its very first step: reward unnecessary escalation and an agent learns to escalate. Hamel Husain’s trial exposes that trap. Kolibri Has to Earn Replacement under the same discipline: refusing unsupported answers helps only if useful answers survive. Better scores can conceal worse service when the evaluator silently decides which mistakes count.

The Grader Chooses the Winner

Agents can now build the test they will optimize against. The useful practice is to separate discovering failures from deciding what counts as success. Otherwise, an automated improvement loop can make a questionable product decision increasingly consistent.

Anthropic’s September 28 eval-design workflow supplies a concrete example. Its public claude-api skill builds evaluations and iteratively changes prompts, settings or code against them. On an internal support benchmark, the company reports improving held-out decision accuracy from 78.6% to 90.5% at about one-fifth the token cost. That held-out slice contained 14 tickets; the search used another 30. These are vendor results on a small task set, not a general productivity estimate.

The mechanism resembles adjusting a thermostat after checking the thermometer. The optimizer changes the system until the measured result improves; the grader determines what improvement means. A grader that mistakes an unnecessary escalation for good service can reward the wrong behavior perfectly. Check the measuring instrument comes before tuning the worker.

Husain and Isaac Flath’s September 30 leasing-assistant trial exposed premature decisions: selecting a failure before inspecting conversations, then requesting approval of grades from aggregate counts. Neither step gave the reviewers enough evidence.

They replaced disconnected Markdown review with an annotation web app. Husain still preferred to hold off on adopting the plugin until its exploration and review workflow improved. This extends independent model review with a practical requirement: a correction belongs beside the conversation that justifies it, so another reviewer can understand the disagreement.

The strongest counterargument is that discovery already worked: Husain rated the one-shot issue finding above other approaches he had tried. The useful division of labor is therefore to let agents propose failures while withholding judgment on the resulting measurement until its cases are inspectable.

This is earlier-week practice evidence, not a fresh release or a measured uplift from the annotation interface. The transferable workflow is narrower: inspect representative failures, correct labels beside their evidence, then allow optimization against the approved criterion. It costs human attention before it saves iteration time.

The payoff is testable: on newly sampled conversations, does the revised grader agree more often with independent human judgments without hiding a recurring failure category?

Kolibri Has to Earn Replacement

Kolibri adds a runnable German-English model to the deployment shortlist, but its own evidence argues against replacing an existing model on specialization alone. The useful question is whether its particular strengths reduce the expensive failures in a real document workflow.

Aleph Alpha released Kolibri’s weights on October 3 under Apache 2.0, with a vLLM serving plugin. Roughly 3.46 billion of its 78 billion parameters are active per token. Lower per-step computation does not make the full model small enough for an ordinary laptop.

The more interesting design choice is deliberate abstention: declining to answer when the evidence does not support a claim. Think of a document assistant asked for a contract’s renewal date. In Aleph Alpha’s training procedure, one version of the document preserves the answer and another removes the necessary evidence. The model must distinguish them. A lucky guess on the incomplete version still fails. The target is evidence-sensitive answering.

That changes the deployment objective. A conventional accuracy score can reward guessing because some guesses happen to be correct. A document service may instead prefer an explicit referral to a human. But referrals have a cost too: a model that declines answerable questions transfers work to the review queue.

Kolibri’s published comparisons do not establish a categorical reliability lead. Aleph Alpha reports a non-hallucination score of 44.0 against Qwen3.6-35B-A3B’s 56.7 on AA-Omniscience, a knowledge test. That is a different question from whether an answer follows from a supplied document. These are vendor measurements, not an independent reproduction. Specialization needs to win on the failure type the application actually faces.

The strongest case for Kolibri is the combination of downloadable weights, bilingual emphasis and a documented serving path. The strongest objection is equally concrete: an existing open model may already handle the relevant exceptions better. The model card also lists roughly 78 GB for the weights alone, making deployment capacity part of the comparison.

The practice consequence is to compare both answerable and deliberately incomplete documents while keeping the downstream escalation rule fixed. This follows the grader story directly: changing which errors receive credit can change the apparent winner without changing either model. Kolibri earns replacement if it lowers unsupported answers at the same rate of correctly completed requests and an acceptable human-review load.

The Contrarian Take

Everyone says: Let agents discover their mistakes, write evaluations and improve themselves; the feedback loop will compound.

Here’s why that’s incomplete: Discovery and grading have different failure modes. A useful observation can become a poor binary test when several distinct problems are bundled together. Kolibri’s refusal training illustrates the other side: an answer can be wrong to give even when its factual content happens to be correct, because the supplied evidence did not justify it. My inference is that the scarce input shifts toward representative cases and explicit error costs as optimization gets cheaper. More iterations cannot reveal a failure category excluded from the score. Human review pays most when it changes what the system is trying to improve.

Under the Radar

  • A million-token ceiling is not the operating recommendation. Kolibri’s card advertises a validated context of 1,048,576 tokens while recommending no more than 262,144 for serving efficiency and complex tasks. Unlike October 1’s Gemini output expansion, this concerns the model’s context capacity. The useful boundary is workload quality and responsiveness at the chosen length, not whether a request fits the advertised maximum. (Model card)
  • One failing grade can conceal four different repairs. Husain’s leasing trial produced one evaluator covering consent, timing, speech during transfer and spoken tool mechanics. Three checks were programmatic; one used a model judge. Separate results would distinguish a deterministic sequencing defect from disputed interpretation. This is a diagnostic refinement to existing review practice, supported by the observed artifact, not evidence of a measured reduction in production failures. (Field report)

Quick Takes

  • Reflection is a future comparison, not a current model pick. Axios reports that Reflection is preparing an open-weight release. That is advance reporting, not a runnable capability result. A useful technical comparison requires accessible weights, license terms, a serving recipe and task results under comparable conditions. Until those exist, the announcement cannot establish whether a deployed coding or research workflow gains anything from switching. (Source)
  • Kolibri runs on Spark, with a substantial startup wait. An October 4 operator report supplies code and raw measurements: roughly 20 generated tokens per second on one DGX Spark, with 18–25-minute cold starts. That makes keeping a service resident a different proposition from loading it for occasional requests. The author’s small test does not establish production reliability, and the standalone Compose configuration was checked syntactically rather than launched. (Source)
  • A probability needs a task-specific cutoff. A September 29 Jev study releases code and responses for 346,009 requests across 37 datasets. Training-set threshold tuning raised its combined precision-and-recall score on unfair-contract-clause detection from 0.50 to 0.75. This adds a different deployment boundary to October 2’s answer-consistency discussion: repeatable probabilities can still produce poor decisions at an arbitrary cutoff. Thresholds need validation on separate examples. (Source)

The Thread

Evaluation criteria can become a hidden source of switching costs. A team chooses a model, builds tests around the failures it notices, and then uses those tests to choose its next model. My inference is that this loop can preserve the incumbent’s strengths while overlooking work it never handled well enough to attract users. Anthropic’s eval-design article explicitly warns that user traffic can skew toward what people already expect to work. Kolibri makes the consequence concrete: a comparison centered on completed answers and one centered on unsupported claims may select different systems. This extends yesterday’s model-and-harness fit with a question about who defines the workload itself. A useful migration evaluation therefore preserves ordinary traffic while adding independently chosen difficult and unanswerable cases. The revealing outcome is whether the model ranking survives that change. If it does not, the organization has learned something about its previous definition of useful work, not just about the challenger.

Prediction Ledger

Weekly Scorecard

  • A non-OpenAI model would show an ARC-AGI-3 harness gap above ten percentage points. Made September 4, medium confidence; due October 4. Correct: ARC Prize lists Gemini 3.8 Flash at 35.00% with its Provider Adapter and 10.37% with its Standard harness at high effort, a 24.63-point difference. This meets the stated threshold; it does not isolate which adapter component caused the gain. (Results)
  • Ollama would ship model-specific thinking controls and defaults in a stable release. Made September 20, medium confidence; due October 4. Pending verification: release-index claims surfaced, but the inspected stable release notes did not establish the exact promised feature. The deadline stays unchanged; an incomplete source check is neither a confirmed success nor evidence of failure.

The tracker still contains 108 overdue AI forecasts awaiting adequate resolution evidence, including the Ollama call. This is a calibration backlog, not 108 implied successes. No new forecast is added this Monday.


Issue date: October 5, 2026. Prepared at 03:49 AM ET.

Tomorrow morning in your inbox.

Subscribe for free. 10-minute read, every weekday.