AI Intelligence

Work Has No Compiler

7 stories · ~7 min read

Work Has No Compiler

Listen

If You Only Read One Thing

Code gave AI agents a luxury most office work cannot offer: an automatic verdict. OpenAI's push to turn Codex into ChatGPT Work gives Agents Leave the Test Suite its scale; a new review finding 187 incompatible measures gives Productivity Has 187 Definitions its warning. As agents spread across work, output expands while proof decays. That is the real adoption ceiling.

Agents Leave the Test Suite

ChatGPT Work is not merely Codex with office plugins. It is an attempt to export the coding agent's loop into jobs where neither the task nor success can be compiled. The model can now act across email, Slack, files, calendars, and browsers. What it loses is software's unusually clear definition of “done.”

OpenAI says ChatGPT Work carries Codex technology into projects that can run for hours. It can create sheets, slides, documents, and web apps from connected systems. OpenAI reported more than five million weekly Codex users, including more than one million people using it outside software development.

An OpenAI-backed study shows how far distribution is from adoption. Codex use reached 98% inside the company, 17% among organizational subscribers, and below 1% among individual subscribers. TechCrunch reports the combined ChatGPT Work and Codex app has 20 million users against more than one billion people who prompt ChatGPT. Its tester consumed 80 million tokens in four days of Work use, an estimated $65 of inference behind a $20 subscription.

Code supplied an outcome contract. Think of it as a machine-readable agreement between an agent and its environment: the patch applies, tests pass, and a reviewer can inspect the diff. A strategy deck has no equivalent compiler. Its effect might arrive months later, after the agent has already received a green check from its own harness.

That missing contract also turns permissions into product design. The TechCrunch test could not grant a cloud drive simple read-only access without repeated errors; the interface eventually asked for full access. OpenAI can add approvals and auto-review, but those controls answer whether an action is allowed. They do not establish whether a sales plan, forecast, or investment memo is correct.

The strongest counterargument is that better models and richer traces will close the gap. OpenAI combines GDPval, a knowledge-work benchmark, with user feedback. Yet code agents improved partly because repositories generated cheap, objective feedback. General work offers delayed, political, or private outcomes. That gives vertical tools such as Harvey and Clay an advantage: they can own the workflow-specific evidence that teaches the agent what success looks like.

The decisive signal is now an outcome metric, not another connector. If OpenAI's next adoption report separates repeat Work use and accepted artifacts from joint Work/Codex accounts, the general-purpose loop is learning to measure its own product. Another combined user count would confirm that distribution is still outrunning proof.

Productivity Has 187 Definitions

AI coding tools are producing more activity than the word “productivity” can safely hold. A newly accepted ASE 2026 review found 187 distinct productivity and productivity-adjacent measures across the literature. A typical study used only five.

The fragmentation is not academic housekeeping. Time saved, commits written, developer satisfaction, defects, releases, and customer use describe different stages of the production chain. The review found positive effects were stronger in perceptions and directional findings than in objective or statistically tested results. “AI makes developers faster” often means the study chose an upstream measure where the tool is designed to look fast.

A separate study of more than 100,000 GitHub developers makes the leak visible. Cumulative adoption of autocomplete, interactive agents, and autonomous agents raised commits by 40%, 140%, and 180%, respectively. The 180% activity lift shrank to 50% at the project level and 30% for actual releases. Across four app marketplaces, new app creation rose moderately, but total usage did not.

The mechanism is a weak link, the slowest complementary stage that caps the entire pipeline. Imagine tripling the speed of an assembly station while inspection and shipping stay fixed. Inventory piles up between them. In software, AI expands code supply while review, architecture, integration, release management, and demand remain human-bound. The study's estimated AI-human substitution elasticity of 0.25 says those inputs are strong complements, not easy replacements.

August 20's field evidence showed the same mechanism closer to the pull request: throughput rose 2.09 times while reviewer load roughly doubled. Today's review changes the interpretation. There is no single productivity number waiting to settle the debate because each number sits at a different point in the chain.

The fair objection is that releases lag adoption and a 30% release gain is still economically meaningful. Correct. The evidence does not show agents are useless; it shows that upstream acceleration is not final output. A field study that produces more than a 60% release lift while holding review load and incident rates flat would weaken the bottleneck thesis. Until then, commit growth is evidence of production pressure, not proof of delivered value.

The Contrarian Take

Everyone says: General-purpose agents win when they connect to more apps and models generate more work.

Here's why that's wrong (or at least incomplete): Every connector expands authority without supplying a definition of correctness. Coding worked as the first agent market because tests, diffs, and releases made output unusually legible. ChatGPT Work is entering domains where effects arrive late, while the productivity literature is already split across 187 proxies. The durable advantage belongs to the system that owns the outcome contract, not the one that merely owns the longest tool list.

Under the Radar

  • Token pricing now has an index before agents have a cost unit. IFX tracks 29 models in a fixed 3:1 input-to-output basket and exposes a local code scanner for model choice, caching, and retries. Its listed prices span $0.06 to $11.25 per million blended tokens. Useful market plumbing, but a token index cannot see failed tool loops or accepted work.

  • StateM turns agent postmortems into executable gates. Its open-source runbook keeps phase, evidence, and legal transitions outside the model's growing transcript. The authors report GPT-5.5 xhigh rising from an 83.1% public reference to 92.1% on Terminal-Bench 2.1. That is not a matched rerun, but it shows why durable procedure can act like a model upgrade.

Quick Takes

  • Kiro prices the harness, but hides the denominator. OpenAI's full GPT-5.6 family is now available in AWS's spec-driven coding agent. Joint testing claims Terra cut cost per successful Terminal-Bench 2.1 task by roughly 82%, the right economic unit. The announcement omits the comparison baseline and matched accuracy, effort, and token budgets, so it proves availability, not a universal saving. (Source)

  • Accuracy is improving faster than reliability. Princeton's Holistic Agent Leaderboard says 24 months of model progress produced only small overall reliability gains, with open-ended tasks barely improving. Larger models can be less consistent because they expose more solution paths. A model that succeeds once and an agent that repeats the success are different products. (Source)

  • Agent security is becoming replayable. A Kaggle competition backed by OpenAI, Google, and IEEE asks entrants to find multi-step paths from untrusted input to unsafe tool actions in a deterministic offline environment. The $50,000 challenge closes entries today and submissions September 1. Replayable attacks turn prompt-injection anecdotes into regression cases. (Source)

The Thread

AI's work surface is expanding faster than its accounting surface. ChatGPT Work can reach more systems, but office outcomes are harder to compile than code. Coding agents can generate more commits, but releases and usage expose the human weak link. IFX counts tokens, StateM persists procedure, and Kiro reports cost per successful task. Each moves one step closer to the missing object: an outcome contract that connects model activity to accepted work.

Predictions

New predictions:

  • I predict: By December 31, at least one of OpenAI, Anthropic, or GitHub will publish a coding-agent productivity study whose primary outcome is releases, deployments, or production incidents rather than commits, pull requests, tokens, or self-reported time saved. (Confidence: medium; Check by: 2026-12-31)

Issue date: August 25, 2026 · Generated: 04:00 AM ET

Tomorrow morning in your inbox.

Subscribe for free. 10-minute read, every weekday.