AI Intelligence

Gemini’s Million Token Bet

7 stories · ~7 min read

Gemini’s Million Token Bet

Listen

If You Only Read One Thing

Google has given Gemini room to generate a million tokens; a research team has made agents better by letting them delete context. Gemini Expands the Workspace follows Argon’s frontier debut. Agents Rewrite Their Working Memory tests the complementary idea: long tasks need control over what stays in view. Capacity and memory management are becoming dimensions of capability, with different costs.

Gemini Expands the Workspace

Gemini 4 Argon puts Google back into serious consideration for difficult engineering and knowledge work. Its most consequential design change is room for much longer reasoning and generation, although the first independent results argue for choosing it by workload rather than declaring a universal winner.

Google’s September 30 announcement raises the output limit from 64,000 to one million tokens. This is output capacity: how much the model can generate, including its extended reasoning trajectory. It is different from the input window that determines how much material the model can read. A bigger desk holds more documents; a longer working session permits more intermediate work before a response must end.

That distinction matters for a migration spanning many files. Additional output headroom can accommodate a longer chain of investigation, implementation and correction. It cannot establish that the investigation stayed useful. The practical question is whether a task previously interrupted by the generation limit now reaches a correct, reviewable result.

The evidence is substantial but uneven. Google reports 77.9% on DeepSWE v1.1, its cited test of extended software-engineering work. Separately, Vals’ evaluation records 57.58% on Terminal-Bench 4.0, a suite of tasks performed through a terminal, placing Argon fifth in that comparison. These are different task sets and scoring systems; subtracting their percentages would tell us nothing. Their different competitive positions do tell us that “good at coding” is too broad a deployment category.

The benchmark comparison does not establish response time or completed-job cost. A model can finish more difficult work while taking longer on ordinary assignments; a better success rate can also conceal more expensive failed attempts. Those measurements need separate evidence. This extends September 22’s model-and-agent fit argument: even a major model advance arrives with a particular pattern of strengths.

The strongest objection is immediate: most people cannot test it. Initial access goes to trusted cyber defenders through Google’s Fairwind program; broader access is promised, beginning with paid API customers and AI Ultra subscribers. That makes Argon a consequential frontier signal today, rather than an available replacement for the daily coding stack.

The decisive next measurement is completed, independently graded repository work at matched time budgets. If Argon’s advantage appears only when competing systems receive less time to finish, the long-output feature has demonstrated additional spending capacity before it has demonstrated superior engineering.

Agents Rewrite Their Working Memory

An agent can improve without a stronger underlying model when it controls which parts of its history remain in view. New research from the University of Washington, Meta and collaborators tests that proposition directly, with results relevant to coding and research agents that repeatedly exhaust their working context.

The familiar workflow lets conversation accumulate until a harness summarizes it. Context Language Models instead expose the live input as an editable file. Think of a working notebook whose obsolete entries can be removed and whose current task state can be revised in place. The model edits that file through ordinary shell operations; the harness synchronizes the changes into its next input. The notebook becomes part of the computation.

The September 29 paper, surfaced September 30, compares context strategies using the same Qwen3.6-27B model and a 32,000-token budget. On BrowseComp-Plus, a demanding research task suite, editable context raises accuracy 11.4% relative to the strongest baseline while using 21.5% less estimated compute. That is a relative improvement, not an 11.4-percentage-point gain. On TerminalBench 2.1, it matches the strongest baseline’s accuracy with 29.5% less estimated compute.

The mechanism goes beyond maintaining a good handoff summary. External notes can preserve evidence while the live conversation still contains stale plans and bulky tool results. Here, the agent can change the material that the next model call actually receives. September 24’s memory analysis examined shorter retrieved records; this experiment changes who controls the working input throughout execution.

The paper links a public implementation, making the approach inspectable beyond a diagram. Its qualitative examples show agents replacing bulky search results with short references and maintaining progress notes in place. These examples illustrate a mechanism; they do not establish a native Claude Code or Codex feature.

The cost claim also needs its boundary. The paper estimates computational operations, accounting for the repeated input processing caused by edits. It does not measure a matching reduction in API invoices, subscription consumption or elapsed time. Aggressive deletion can also remove something needed later; unrestricted editing should not be mistaken for evidence of safe retention.

The practical refinement is selective replacement of live context, beyond the existing habit of moving a long session to a fresh manager. It pays off only when matched tasks retain their acceptance rate while reducing measured end-to-end cost, including reconstructing anything the agent deleted prematurely.

The Contrarian Take

Everyone says: Bigger context and longer reasoning will make long-running agents more reliable.

Here’s why that’s incomplete: The editable-context experiment improves research accuracy while reducing estimated compute, so additional retained text is plainly not the only route to better work. Argon’s larger output allowance addresses a different constraint: generation can continue for longer, but the extra room does not determine which intermediate work deserves to survive. Treating both developments as “more memory” conceals the design choice between preserving evidence and repeatedly presenting it to the model. The emerging capability is the ability to revise a working state while retaining an auditable history elsewhere; today’s experiments support the first half more clearly than the complete production design.

Under the Radar

  • A correct embedding shape can hide the wrong operation. Perplexity’s new model card requires encode_queries for search queries and encode for document chunks. Both produce vectors, but the wrong method silently degrades retrieval because query and document training prefixes differ. A pipeline can pass its type checks while giving the search model the wrong kind of input; a small query-to-document regression set tests behavior that schema validation misses.

  • Generated software can keep private inputs out of subsequent model calls. Simon Willison’s September 29 Photo Scrubber runs face detection and image redaction locally in the browser, then rebuilds an export without the original metadata. The useful practice is generating a reusable local processor instead of sending every photograph to a hosted model. Detection can miss faces, so the tool includes manual regions and an explicit review step; no anonymization guarantee or measured time saving follows.

Quick Takes

  • Perplexity changes what retrieval learns to keep. Its September 30 contextual-embedding preview encodes document chunks together, preserving surrounding meaning in each searchable vector. For research agents, that creates a route to finding passages whose meaning depends on evidence elsewhere in the document. Retrieval can fail when an isolated paragraph loses the context needed to interpret it. The open model remains a preview; later embeddings may be incompatible with today’s index. (Source)

  • Cloudflare makes short-lived agent workspaces easier to use. Its September 30 Containers rebuild allows image and machine-size selection at runtime and adds filesystem snapshots in public beta. Cloudflare cites ComputeSDK measurements showing median startup falling from just over four seconds to 648 milliseconds. Lower startup overhead makes a fresh environment for each task less expensive in waiting time; it does not prove faster complete tasks or that a filesystem snapshot preserves a running process. (Source)

  • Sonnet 4.5 now has a retirement date. Anthropic’s September 30 notice schedules its Claude API retirement for November 30 and names Sonnet 5.5 as the replacement. This adds a concrete migration deadline to the capability release covered September 29. A pinned model identifier preserves reproducibility only while the provider serves it; acceptance examples and recorded behavior must survive independently of that endpoint. The notice concerns the Claude API, not a universal shutdown date across every reseller. (Source)

The Thread

Longer reasoning makes the quality of forgetting more consequential. A migration agent can spend its extra generation budget following an obsolete diagnosis, or delete that diagnosis and lose the evidence that would have exposed its mistake. Argon expands the possible trajectory; editable context changes which earlier steps remain available to guide it.

September 24’s storage-versus-input argument separated an archive from a working view. The new dependency is that the worker can now revise that view during execution. Keeping the original transcript permits later inspection, but does not make a mistaken deletion visible to the agent at the moment it matters.

My inference is that context editing needs a recovery test alongside its efficiency test. Replaying the same task with an omitted constraint restored would reveal whether the saving depended on forgetting something important. The paper’s aggregate accuracy and compute results support the approach; they do not establish that every deletion is harmless. As reasoning runs grow longer, a useful memory system must make consequential omissions detectable before the final answer, while recovering enough evidence to correct them.

Predictions

  • I predict: Google will open Gemini 4 Argon to at least one paid developer or AI Ultra cohort outside the initial Fairwind and trusted-testing programs by October 31. The stated rollout path supports the forecast; safeguards or capacity could delay it. (Confidence: medium; Check by: 2026-10-31)

October 1, 2026 · 03:36 AM ET

Tomorrow morning in your inbox.

Subscribe for free. 10-minute read, every weekday.