Skip to content
HN On Hacker News ↗

I stress-tested Meta Muse until its agent control plane started timing out

▲ 8 points • 9 comments • by mpkc • 4w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is AI.

100 %

AI likelihood · overall

AI
0% human-written 100% AI-generated
SEGMENTS · HUMAN 0 of 1
SEGMENTS · AI 1 of 1
WORD COUNT 1,572
PEAK AI % 100% · §1
Analyzed
Sep 14
backend: pangram/v3.3
Segments scanned
1 windows
avg 1572 words each
Distribution
0 / 100%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 1,572 words · 1 segments analyzed

Human AI-generated
§1 AI · 100%

Meta Muse is Meta’s personal AI agent, launched on September 8, 2026. Rather than only answering questions, it is designed to carry out tasks on a user’s behalf: it has its own browser, can keep working after the app is closed, and runs in a dedicated Muse Secure VM. In presentation it resembles Grok Bot — a personified agent controlled through conversation — but that is an interface-level analogy, not an assumption of shared architecture. This article goes one layer lower and examines a narrow part of the runtime: subagent fan-out, the durable state it leaves behind, and the spawn path under load. I ran these tests in my own Muse session, using only interfaces exposed to that session: subagent.spawn, a shell in the assigned environment, and a bounded read-only interface to durable diagnostic and trace data. I did not attempt to access other users, tenants, or data outside the environment assigned to me, and I did not bypass access controls. I am not presenting this as a Meta-authorized security assessment, and access alone is not evidence that every load experiment was separately authorized. This is a black-box/reverse-engineering write-up from the perspective of the access granted to my session. Post-publication update — September 13, 2026. I clarified the access scope, added independent architectural context from Rohan Adwankar’s analysis, separated STAGGERED-80 from the burst-style runs, and described two controls that would better isolate cadence, topology, and concurrency. The experimental data, published CSVs, and figures were not changed. At 06:45:32 UTC I asked a chat session to spawn 120 subagents at once. Each one had a deliberately trivial job: run sleep 30 in a shell and report a single line back. Thirty-three of those calls created an agent. Eighty-seven failed with the same database error. The aggregated answer never arrived, and the interface eventually showed an Error state. In the durable trace, every one of the 33 created agents reached a terminal completed record - 32 of them with a confirmed workload completion, and one still unresolved - while the record of the parent still said running. What follows is a black-box investigation of that gap, built entirely from records the runtime wrote to PostgreSQL as it worked: an agent registry, a spawn ledger, per-worker progress tables and a context-item store. The load tests were not re-run to write this article. Everything here was reconstructed from stored state, with one documented exception I will come back to. What Muse looked like from the outside From a user’s seat this was an ordinary chat session. The persona text was plain; the tool list was not. Among the tools the session could call were subagent.spawn, subagent.close and subagent.resume, plus a read-only interface to a PostgreSQL database. That database was not incidental - it held the runtime’s own bookkeeping. Tables such as agent.agents, agent.subagent_spawns and agent.subagent_progress_tool_events recorded which agents existed, who spawned them, which tools they ran and when they finished. That is how the multi-agent structure became visible at all. Not through the interface, which presented one conversation, but through the records the runtime kept for itself. Two labels need care before anything else in this article. The first is the model string. Every agent row I read - the root, the coordinator, all 33 workers of the largest burst - carried the same value: ipnext/avocado-5.16-v4 That string is observed. What it means is not. It could be an internal model build, a routing alias, or something else entirely; the trace does not say, and I am not going to guess. I am publishing it verbatim because it is one of the few hard identifiers the durable trace provides. The second is the runtime’s description of itself: Muse Spark 1.3, from Meta’s Muse family. That comes from the runtime’s own context rather than the trace, so it is self-reported, not verified. Everything else in this article is anchored in records. Independent architectural context After publication I found an independent teardown of a contemporary Muse instance by Rohan Adwankar. In his environment, PostgreSQL ran inside the per-user VM over a local Unix socket, and the harness binary contained both avocado-5.16-v4 and ipnext/... paths. That fits several of my observations and gives them useful architectural context, but it does not change the boundary of the evidence: my data still does not identify the specific lock, table, row, index, query, or transaction responsible for the timeouts. I still treat ipnext/avocado-5.16-v4 as an observed model identifier. Adwankar interprets ipnext as Meta’s internal transport/gateway and avocado as an internal model family; directly mapping avocado to Muse Spark 1.3 remains an inference, not something established by my trace. Likewise, my own probe established KVM visibility, while Adwankar identified Cloud Hypervisor running on KVM in his instance. I did not independently establish the specific VMM used by mine. External source: Rohan Adwankar, “What’s in a Muse?” Method: counting overlapping agents The experiments used one workload, four configurations and no retries. Three of them - PROBE-40, BURST-80 and BURST-120 - were burst-style runs issued from the root agent: every attempt in a configuration went out in a single turn. The fourth, STAGGERED-80, was deliberately spread over time and used a coordinator agent to spawn its workers. No failed call was repeated. Concurrency here has a narrow, deliberate definition. An agent counts as active from its first tool call until its terminal record: active(t) := first_tool_at <= t < finished_at Concurrency over time is a sweep line over those intervals, with ends processed before starts when timestamps tie. The timestamps have one-second resolution, so the sweep is deterministic for the recorded data but cannot recover event ordering within the same second. Three consequences apply to every number below: This measures agent activity, not inference. Thirty-three overlapping activity windows are not thirty-three simultaneous model calls. Nothing here measures the inference backend. A peak is an observation, not a limit. The tests never exceeded 120 simultaneous attempts, so they cannot establish a concurrency cap - or rule one out. A missing record is evidence. Failed spawn calls leave nothing behind in the agent registry. That asymmetry shaped how failures had to be verified, and it is why the failure counts were the hardest numbers to pin down. The first probe: 40 calls The first experiment was a calibration run: 40 spawn calls issued at 06:10:36 UTC. It created 39 agents, admitted within two seconds of each other. One call failed, and that failure is worth a closer look, because it was recovered from the durable tool trace rather than taken from the chat: the call went out at 06:10:44 UTC, the error came back at 06:10:56 UTC, and no child agent was created. The peak observed concurrency was 39. That figure is an archival recomputation from the session’s trace table, not a live measurement, and the published summary says so. One more detail from that reconstruction carries weaker provenance than anything else in this section. In the session-built table, eleven of the thirty-nine created agents finished with a background-processing status message instead of the requested DONE line. Because it survives only in a hand-built table, I report it with that caveat and draw no conclusions from it - except that “terminal status” and “the workload finished” are not the same statement. That distinction matters much more later. Eighty at once, and eighty spread out The next two experiments look like a clean A/B test. They are not. Burst-style runs Configuration Attempts Created Failed Failure rate Peak observed concurrency PROBE-40 (burst) 40 39 1 2.5% 39 BURST-80 (burst) 80 75 5 6.25% 72 BURST-120 (burst) 120 33 87 72.5% 33 STAGGERED-80 — separate configuration I show STAGGERED-80 separately because this run changed both cadence and topology. Its 0% failure rate is therefore not a fourth point in the same series as 2.5%, 6.25%, and 72.5%. Configuration Attempts Created Failed Failure rate Peak observed concurrency STAGGERED-80 (spread) 80 80 0 0% 38 BURST-80 issued 80 spawn calls at once, at 06:30:43 UTC. Seventy-five agents were created - all of them direct children of the root agent - and five calls failed. All five returned the identical database lock timeout, each about 58 seconds after the call was placed. The peak observed concurrency was 72, recomputed after the run from the analysis input for that phase. STAGGERED-80 started at 06:32:03 UTC and issued the same number of calls, spread out. It created all 80, with zero failures. It is also the experiment that refuses to be a clean control, because two things changed at once. First the cadence: the intended spacing was 100-200 milliseconds, but the measured mean admission gap was 1.1266 seconds, roughly ten times wider than planned. Then the topology: the root spawned one coordinator at depth 1, and the coordinator spawned the 80 workers at depth 2. The burst experiments spawned workers directly from the root. The staggered run also peaked lower - 38 active workers against 72 in the comparable burst - so the two runs differ on more than their failure counts. Observed active-agent concurrency over time. Activity is defined as first_tool_at <= t < finished_at, with finishes processed before starts on timestamp ties; timestamps have one-second resolution. Each configuration was run once. STAGGERED-80 also changed topology, so this is not a cadence-only comparison; the curves show agent activity, not inference concurrency. Source: BURST-80 aggregate reconstructed from the complete archived trace; STAGGERED-80 independently recomputable from the published worker rows. Three admission shapes, three observed outcomes.