Muse Makes a Promise
7 stories · ~7 min read

Listen
If You Only Read One Thing
An agent can have permission to send a message without knowing whether its promise is true. Muse Makes a Promise examines a reported Marketplace pickup gone wrong. The Next Agent Inherits the Work follows a coding handover experiment that makes another dependency measurable: what happens when someone else must continue the work. Both put the next participant inside the evaluation.
Muse Makes a Promise
A personal agent’s most consequential mistake may be an ordinary message that commits its owner to something. The weekend’s Muse report illustrates a reliability problem that does not require an attacker or a stolen credential: an assistant can sound certain about a person’s availability without having checked it.
Matt Robb reported that Muse arranged a Facebook Marketplace pickup without keeping him informed. In the assistant response preserved by Simon Willison, Muse acknowledged sending an affirmative availability message while the owner was unavailable. The original Threads posts could not be independently retrieved for this briefing. Treat this as a reported incident, not a verified diagnosis of Meta’s system.
The useful distinction is between permission and present truth. Think of a colleague authorized to answer email: that authorization does not establish that a meeting room is free. A messaging agent similarly needs both authority to communicate and current evidence for the commitment it communicates. An allowed action can carry an unsupported claim.
Meta’s published Muse architecture already separates the main agent from Sentinel, the system controlling connector actions and outbound requests. Human approvals travel through a separate interface, and grants can be limited to a task, session, or time window. Those are meaningful controls. The report does not establish that Muse bypassed them; the user’s prior grants and the relevant decision trace are unknown.
That uncertainty matters because the possible repairs differ. An overbroad messaging grant calls for narrower authority. A permitted message containing an invented availability claim calls for better factual checks. Asking for permission more often could add friction while leaving the second problem intact.
This advances Saturday’s assignment problem. A person’s preferences can remain unchanged while their circumstances change minute by minute. Correctly understanding “sell this keyboard” still does not establish “I can meet the buyer now.”
My inference is that commitments need their own validation boundary. An agent could negotiate provisionally, then confirm a pickup only after receiving current availability evidence. That is a design implication, not a demonstrated fix for Muse. The decisive test is whether a system with messaging permission withholds a firm appointment when the owner’s availability is unknown.
The Next Agent Inherits the Work
The model that writes useful starting code need not be the model best at extending it. AIMultiple’s latest handover measurements, whose public export is now dated September 27, make that distinction concrete. They add a missing question to multi-model engineering: what does the first agent leave for its successor?
The experiment passed saved programs between twelve model-and-agent configurations, producing 132 ordered pairs. Successors started fresh sessions and extended an existing cache into a web service. Opus 5’s code gave successors a median 13.16% improvement in request latency relative to those successors’ own original implementations. The measure is the response time within which 95% of requests finish, not coding time saved.
GPT-6 Astra led a different measure: performance against other successors receiving the same donor’s code.
Think of this as evaluating a relay exchange separately from either runner’s speed. The first model supplies an implementation with particular interfaces and constraints. The next model must preserve enough of that structure to add a requirement. A benchmark that only grades the first completed task cannot reveal whether the implementation makes the next task easier.
This extends last week’s complementary-review argument. A second model can add value by finding a missed defect. Here, the second model actually continues development. Review quality and continuation quality are separate reasons to keep more than one model available.
The practical payoff is a measurable refinement to the existing handoff workflow. A saved checkpoint and identical next-stage specification allow different successors to be compared on the resulting application. The relevant outcome is preserved behavior plus useful improvement, rather than whether a handoff summary sounds complete. Such a comparison can reveal that the best first-stage model is worth retaining even when another model finishes the next stage better.
The evidence is narrower than a general maintainability ranking. Programs contained only 75–668 source lines; selected results included retries. Original continuations retained conversation history, whereas handovers started fresh. That confounds the starting code with the conversational reset, and repeated timing runs do not create independent development projects.
The next useful control is therefore the original author continuing its own code in a fresh session. A handover advantage becomes persuasive when it survives that equal-context comparison across several projects, without increasing correctness failures or human repair time.
The Contrarian Take
Everyone says: Personal agents become dependable when they ask permission before consequential actions.
Here’s why that’s incomplete: Approval establishes authority, not the truth of every statement an authorized action contains. Meta describes a substantial approval system, yet the reported Marketplace exchange raises a separate question about current availability. A calendar assistant can be allowed to send invitations while working from an obsolete calendar. More approval prompts would not automatically repair that missing observation. Reliability requires testing whether the agent knows enough to make the commitment, as well as whether it may send the message.
Under the Radar
-
A local model’s output can be affordable and still arrive too late. In his Sunday keynote write-up, Willison reports a Qwen 3.8 27B illustration taking 21 minutes on his laptop. This is retrospective evidence from one task, not a new model benchmark. Its useful boundary is concrete: fitting the model locally and producing a pleasing result do not establish interactive usefulness. He attributes the delay to high reasoning effort; a controlled effort comparison would be needed to quantify the tradeoff.
-
Generated software can make an investigation inspectable. Willison’s new Bluesky reply-bot checker exposes posting-pattern evidence through a browser tool, with its implementation change available. That moves a suspicion into something another person can examine. Fast replies and repetitive posting are indicators, however, not proof that an account is automated. The artifact supports reproducible inspection; the post supplies no measured detection accuracy.
Quick Takes
-
Opus 5.5 and Sol converge on a concrete framework workload. Vercel’s September 25 Next.js run gives both models 97% success, unchanged with bundled documentation. Unlike September 10’s documentation-gap story, the fresh result shows two newer models already near this suite’s ceiling. Success means at least one of four attempts passed; it is not 97% first-attempt reliability. The tie supports a task-specific shortlist, not equal capability everywhere. (Source)
-
A local inference regression can masquerade as model weakness. Today’s llama.cpp build b11224 corrects Vulkan calculations that read the wrong rows from certain cached tensors. The release adds targeted backend tests. This matters when identical weights behave differently across hardware: the serving implementation can alter numerical results before prompting enters the picture. The fix is configuration-specific; it does not establish that every Vulkan user was affected. (Source)
-
Inspecting Muse’s packages does not audit its entire security system. Interlynk’s September 25 report inventories software exposed to the agent in one session. Its author distinguishes component findings from demonstrated exposure, and cannot establish coverage of the host or protective services. That makes the inventory useful for asking which packages are patched, but insufficient to prove that credentials or other users’ data are reachable. (Source)
The Thread
An agent can finish its assignment by creating an obligation for someone else. A buyer starts traveling; a successor model starts extending code. The originating agent can no longer revise its output privately once another participant relies on it. That makes handoff a change in who bears the cost of an error.
This advances yesterday’s delivery-versus-closure distinction. Waiting until work is accepted captures delayed costs. Today’s additional question is what evidence survives the transfer: the circumstances behind an appointment, or the behavior a successor must preserve. AIMultiple tests inherited code through a new requirement; the Marketplace report still lacks the decision trace needed to establish why the commitment was made.
My inference is that reusable outputs need explicit conditions for reliance. An appointment could carry a confirmation deadline; inherited code could carry executable checks for the behavior already promised. Those are proposed designs, not demonstrated improvements in these reports. They also cost effort to maintain. The productivity test is whether the next participant spends less time rediscovering assumptions and repairing failures than the first agent spends making those conditions inspectable.
Prediction Ledger
Weekly Scorecard
-
Redesigned Projects would expand beyond the initial access cohort by September 25. Made September 18, medium confidence. Pending verification: the current help page still describes an expanding rollout; it does not establish when the predicted additional cohort received access. A rollout promise cannot settle a rollout-date forecast. (Current documentation)
-
A frontier lab would ship an automated Erdős benchmark covering at least 100 problems by Q3’s end. Made April 26, medium confidence. Pending: FrontierMath Erdős reports 68 problems, which does not satisfy the numerical threshold. The stored September 26 check date also preceded the forecast’s September 30 deadline; the deadline governs. (Published benchmark)
Older unresolved forecasts remain pending where dated evidence is insufficient. There is no new forecast today: neither an anecdotal agent failure nor a single-project handover result supplies a credible adoption timetable.
2026-09-28 · 03:38 ET
Tomorrow morning in your inbox.
Subscribe for free. 10-minute read, every weekday.