AI Intelligence

Discovery Outruns Human Review

7 stories · ~7 min read

Discovery Outruns Human Review

Listen

If You Only Read One Thing

Finding more can make finishing harder. Anthropic’s OSS Scanner sends unreviewed vulnerabilities straight to willing maintainers, while Exa’s ATLAS exposes how much research agents still leave undiscovered. The practical distinction is between checking what an agent returned and establishing what remains unresolved. Better discovery changes both workloads, but a convincing pile of results cannot establish that either job is complete.

Claude Makes Findings Testable

Security-agent findings become more useful when they arrive as executable evidence. Anthropic’s October 8 OSS Scanner launch packages that evidence for eligible open-source maintainers. Today’s News briefing covers the transfer of triage work; the technical question is what makes an unreviewed finding economical to verify.

The company reports more than 29,000 candidate vulnerabilities from six months of scanning, with approximately 6,000 manually reviewed. Its launch account describes an expert audit of 97 high- or critical-severity findings across 48 projects. Eighty-five met its disclosure standard; eleven were real but duplicated known issues or other findings; one was invalid. That is evidence of useful discovery, not a measured reduction in deployed vulnerabilities.

The important output is a reproducer: a small, self-contained test that demonstrates the alleged bug. Think of the difference between a reviewer saying a door might be unlocked and supplying a repeatable check of that particular lock. A second model can dispute the explanation. An executable test lets the maintainer inspect the behavior. This is evidence before agreement, extending independent model review with something the reviewer can actually run.

The service packages explanations and, where available, candidate patches alongside that evidence. Anthropic’s enrollment instructions also let maintainers supply a threat model: which inputs are hostile, which behavior is outside scope, and how severity should be judged. A technically reproducible behavior can still be irrelevant to a project’s security promises. The test establishes what happens; the threat model establishes why it matters.

The strongest objection is workload transfer. Removing the vendor’s human review does not remove the work; it hands that work to the project. The company retains its human-verified disclosure route for maintainers who cannot absorb raw reports. That distinction makes the fast track more credible than a universal claim that every project should receive more findings.

My inference is that security-agent quality will increasingly depend on the cost of adjudicating each report. A plausible paragraph creates reading work. A reproducer, relevant scope and usable patch can turn that work into a bounded engineering task. None proves that the patch preserves the rest of the program.

The adoption test is whether verified fixes per maintainer-hour increase without a growing queue of unresolved high-severity reports. A rising count of discoveries alone cannot pass it.

Research Needs the Missing Rows

A research agent can be right about every company it lists and still fail the assignment. Exhaustive searches require evidence about the companies it omitted. Exa’s October 8 ATLAS preview makes that failure visible across 547 tasks that ask agents to discover qualifying entities and fill in their attributes.

The vendor reports that even its most expensive tested agents miss roughly a third of the reference results. Its headline measure combines accuracy and completeness, counting a row as correct only when every required cell is correct. No tested configuration with a median cost below $1 per task exceeded 0.5 on that combined measure. That does not mean those systems answered exactly half the rows correctly; the score also penalizes incorrect returned rows.

The distinction matters for actual research work. Imagine a supplier comparison in which every quoted price is correct, but the agent never finds the supplier offering the necessary certification. Checking citations would approve the visible evidence. The procurement decision would still be based on an incomplete field. Missing rows change decisions even when none of the printed statements is false.

This extends October 4’s search-provider comparison beyond whether adding search improves answers. The practical refinement is an evaluation with two separate questions: did the agent find the required set, and did it describe the found members correctly? A second reviewer given only the finished report cannot reliably answer the first question. The review needs an independently assembled reference set or a separate discovery pass whose disagreements are investigated.

Exa also reports a useful intervention: hiding the top seven of ten results roughly halved ATLAS performance for GPT-5.6 Luna using Exa search. That supports treating retrieval quality as a cause of failure, rather than automatically buying more reasoning from the same model. It does not establish that any particular search vendor will win on a different workload.

There are substantial limits. Exa sells search and constructed the evaluation; tasks, answer tables and grader are promised for the coming weeks, not available for independent reproduction at launch. Its reference sets can themselves omit valid answers. The result supports a more demanding acceptance test, not an independently settled provider ranking.

For research-agent deployment, the revealing comparison is omission rate at a fixed total budget, measured separately from errors in returned facts. A cheaper run earns replacement only if the missing results do not change the decision.

The Contrarian Take

Everyone says: Once AI finds real bugs reliably, removing human triage is an obvious efficiency gain.

Here’s why that’s incomplete: Anthropic’s pilot audit found eleven genuine-but-duplicate findings and only one invalid finding among the twelve reports that failed its disclosure standard. Truth is therefore not the same as additional useful work. Two accurate reports can send different maintainers toward the same repair, consuming scarce attention without removing another vulnerability. My inference is that deduplication belongs before assignment, with a shared issue identity that survives repeated scans. Otherwise improving factual accuracy can still leave the repair process less efficient.

Under the Radar

  • An offline audit still has an online preparation stage. OSS Scanner’s configuration guide says the container image is built with network access, while the subsequent audit runs without it. That makes the dependency preparation stage a separate trust boundary: hermetic execution constrains the running agent, but does not by itself establish the provenance of what was installed beforehand.

  • A text decision score does not validate image or audio decisions. Liquid AI’s open d1 release reports its text Decision Index results but withholds the private vision split; it also says dedicated audio decision benchmarks remain an open problem. A local multimodal classifier may be runnable before its claimed usefulness on a particular input type has comparable public evidence. Text routing and visual inspection need separate acceptance sets.

Quick Takes

  • Grounding changes the model shortlist. Artificial Analysis’s October 8 Harvey LAB-AA update checks deliverables against source documents. Muse Spark 1.3 falls from 26.7% all-pass to 8.9% after material hallucinations disqualify results; GPT-6 Astra moves from 8.9% to 8.6%. These are legal-document tasks with model judges, not universal rankings. The transferable result is that satisfying a rubric and preserving source facts can favor different models. (Source)

  • Local decisions gain downloadable weights. Liquid AI’s October 7 d1-3B release returns probabilities over defined choices instead of generating prose, with llama.cpp support. Its reported RTX 4090 latency rises from 8 milliseconds for one short question to 102 milliseconds for a 3,400-token state. The deployment opportunity is bounded routing or classification; the fastest headline timing does not describe long-context decisions or concurrent application traffic. (Source)

  • A token counter’s default is part of the experiment. Simon Willison’s ttok 1.0 changes its default from the older GPT-4 tokenizer to o200k_base and exposes model-prefix handling. An unchanged file can therefore produce a different count after a tool upgrade. Local estimates need a recorded tokenizer and tool version; the change is not evidence that the serving model or its billing changed. (Source)

The Thread

Agent output makes some work visible while concealing other work. A scanner turns a possible vulnerability into a report someone now owns. A search agent leaves an omitted supplier invisible, so nobody is assigned to investigate it.

My inference is that organizations will overinvest in clearing visible queues while underinvesting in discovering missing obligations. This differs from October 7’s maintenance-debt argument: the issue is which unfinished work enters the accounting at all.

A security audit needs checks beyond the reported bugs, just as a research report needs searches outside its returned list. Otherwise a faster completion dashboard can coexist with more unresolved exposure. The useful management distinction is between work closed and coverage established; either number alone can improve while the other deteriorates.

Predictions

  • I predict: By November 9, Exa will publicly release at least one ATLAS task set with corresponding answer tables and a runnable grader. Its stated release plan supplies the evidence; delivery of all three artifacts is the test. (Confidence: medium; Check by: 2026-11-09)

Coming Next Week

The next comparison worth developing is whether reproducible agent findings reduce human review time once duplicates, severity disputes and patch regressions are included. Discovery counts leave that question unanswered.


Issue date: 2026-10-09. Generated: 2026-10-09 07:25 UTC.

Tomorrow morning in your inbox.

Subscribe for free. 10-minute read, every weekday.