Skip to content
HN On Hacker News ↗

When the machine boots but the reply disappears · Mainbrella

▲ 4 points • 3 comments • by cs1996 • 2d ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is AI.

100 %

AI likelihood · overall

AI
0% human-written 100% AI-generated
SEGMENTS · HUMAN 0 of 1
SEGMENTS · AI 1 of 1
WORD COUNT 1,496
PEAK AI % 100% · §1
Analyzed
Oct 7
backend: pangram/v3.3
Segments scanned
1 windows
avg 1496 words each
Distribution
0 / 100%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 1,496 words · 1 segments analyzed

Human AI-generated
§1 AI · 100%

← Engineering blog October 7, 2026 · Mainbrella Engineering You ask a server to create a Linux machine. The machine boots. The reply disappears. From the client’s chair, this looks much like a request that never reached the server. Retrying is reasonable. Starting a second machine would make the recovery more expensive than the failure. This is a useful way into Mainbrella’s backend. Follow one machine far enough and the same problem keeps returning: a command can run without its caller seeing the result; a private request can outlive a network membership; a disk capture can succeed without leaving us the handle needed to restore it. Each boundary needs a decision about what a retry is allowed to do. Let’s follow an illustrative job that runs analysis.py, reads data from another container, and writes metrics.csv. Its machine occupies slot c17. The filenames and slot are examples; the mechanisms below come from the open-source backend. We’ll start before Linux boots, because that is where ownership and budget have to be decided. One place decides who gets the last slot Two create requests can arrive while an account has one slot left. Reading a counter, booting a machine, and updating the counter afterward gives both requests a chance to win. By the time we notice, the expensive part has already happened. We give each account a Cloudflare Durable Object: a persistent coordinator with its own storage. The public API resolves the authenticated user and paid entitlement, then addresses that coordinator as account:<userId>. Each machine slot has a separate Durable Object. For c17, its name is user:<userId>:slot:17. The first slot retains the older name user:<userId> and the public ID small; that ID does not select the machine’s size. The account coordinator serializes admission. A request must reserve its slot, monthly start, and compute allowance before it can ask the runtime to boot. The next request therefore sees occupied capacity even while the first machine is still starting. The account controller implements the queue explicitly with a promise tail; an await inside an admission decision doesn’t let the next decision slip past it. Holding this lock until every machine was ready would also serialize all their boot times. We release it after the durable reservation and provision outside it. One account makes the decisions in order; its runtimes can boot together. The short shared decision precedes the longer independent work. Time is schematic: the overlapping bars explain concurrency, not measured startup latency. There’s a particularly useful test for this distinction. It sends six concurrent requests on a Builder account, which has five slots, and holds every runtime’s readiness check behind a gate. Five machines reach boot before that gate opens. The sixth request receives a conflict, and the account records five starts. It checks both halves of the design: exclusive admission and parallel provisioning. Ownership is established before this coordination begins. The authentication adapter accepts an API key or a login session. An explicitly invalid Bearer credential fails authentication even if a valid browser cookie accompanies it. The internal request builder supplies the account identity and entitlement itself. Letting the caller choose an x-mainbrella-user header would undo the account boundary we just built. A receipt exists before the machine does For our analysis.py job, the client supplies an Idempotency-Key with POST /containers. That key means “this particular attempt to create a machine.” The account records which slot and reservation belong to it, together with a fingerprint of the requested configuration. Changing the image, size, or other fingerprinted selection while reusing the key produces a conflict. The important write is small. Here is the relevant part of container-account-core.js, with the surrounding validation omitted: const reservationId = ++state.nextReservationId; state.reservations[slot] = reservationId; const creation = { id: crypto.randomUUID(), slot, reservationId, fingerprint, expiresAt: this.now() + CREATION_RETENTION_MS, }; await this.ctx.storage.put({ [KEY]: state, [CREATION_PREFIX + idempotencyKey]: creation, }); That multi-key write commits the charged reservation and creation receipt atomically. We don’t want a receipt for a slot we never reserved, or a charged slot that a keyed retry cannot find. Only afterward does the controller dispatch the boot. Now lose the reply. A matching retry finds the receipt before attempting fresh admission. While the reservation is pending, it can return starting. Pending reservations have a ninety-second reconciliation window; after that, the coordinator asks the runtime what actually exists. If it finds the running machine, it returns that machine without charging another start. The retry never dispatches the ambiguous reservation a second time. The receipt lasts twenty-four hours. If the machine has stopped or its slot has been reused, that same retained key returns creation_no_longer_running. It doesn’t quietly launch a replacement. A client that wants a replacement makes a new creation attempt with a new key. This is also why the client should save its key and request before sending them: server-side deduplication is little help if the client forgets the identity it needs to ask about. What establishes readiness? The runtime controller starts the selected image with sleep infinity as its entrypoint, then executes uname -a. The command must exit successfully within sixty seconds. This establishes that the guest can execute a command. Our Python analysis still needs its own dependencies and checks; an answering kernel cannot certify a financial calculation. Yesterday’s message can arrive at tomorrow’s machine Suppose the dispatch for our first boot is delayed. Meanwhile, the account cancels the reservation, releases c17, and assigns that slot to a replacement. The old dispatch finally arrives. Its target is still the same runtime object. Looking up the slot by name cannot tell us whether this message is entitled to start anything. Each admitted start receives an increasing reservation number. At the runtime, we persist two high-water marks: the newest accepted start and the newest cancellation. A boot at or below either mark is rejected. Cancellation writes its fence even when there is no running guest to destroy, so a later arrival cannot resurrect canceled work. Follow the dashed diagonal: the first boot arrives after the replacement. The cancellation fence and accepted reservation make its age visible to the runtime. The reverse race matters too. A delayed cleanup for reservation 41 must not stop the machine from reservation 42. The runtime’s DELETE path advances the cancellation mark but leaves the guest alone when the cancellation is older than the accepted start. Back at the account, the completion of a boot must still match the current slot reservation before it can clear pending state. An old successful reply has no authority over a replacement either. Reservation numbers protect the internal lifecycle messages. Public commands and file requests identify a running generation with { id, createdAt }. The slot ID is reusable; the generation is not. Despite its timestamp-shaped representation, createdAt is computed as max(now, previousCreatedAt + 1). Recreating a machine in the same millisecond, or moving the wall clock backward, still gives the new guest a different identity. Keep both values from the returned running container. A cleanup request for just c17 cannot express which lifetime you mean. The lifecycle tests deliberately stop and reuse a slot before releasing an old boot dispatch, including without advancing the clock. This is a more revealing check than another successful hello-world launch. The hard deadline is part of admission Our machine also needs permission to keep spending compute. A monthly counter checked only at launch would let many simultaneous machines consume the same remaining allowance. We reserve the runtime they may use before any of them starts. Machine sizes have weights in plan-policy.js: Lite uses one compute unit, Medium uses ten, and XL uses twenty-eight. The account reserves unit-milliseconds. Its lease ends at the earliest of four boundaries: hard deadline = min( start time + plan session limit, paid access expiration, next UTC month boundary, start time + remaining unit-ms / machine weight ) For an arithmetic example, reserving a Medium machine for one hour commits ten compute-unit-hours. Confirm a stop after five minutes and the consumed amount is 10 × 5 / 60, about 0.833 compute-unit-hours; the unused reservation is released. A failed stop or unreadable runtime keeps its reservation. Treating “couldn’t contact it” as “it must be free” would allow the account to spend that allowance twice. A failed admitted launch still consumes its monthly start; runtime settlement is a separate calculation. The runtime persists this hard deadline and an idle deadline. Real activity can move the idle deadline, capped by the hard deadline. Status polling doesn’t. A Durable Object alarm enforces expiration even when the client has gone away. Plan changes can shorten an existing lifetime, but cannot extend its original hard deadline. Stopping work should remain possible when the billing lookup is unavailable. The public DELETE path authenticates ownership without requiring a fresh billing resolution; the coordinator uses its saved entitlement for cleanup when it remains valid. Likewise, a newer unpaid observation beats an older paid one. Otherwise a delayed check could reauthorize a machine we had already revoked. A disconnected viewer shouldn’t own a process