How to Build an AI Software Factory: Agents That Open, Review, and Merge PRs
Pangram verdict · v3.3
We believe this text is mainly AI, with some human-written content.
AI likelihood · overall
AIArticle text · 793 words · 2 segments analyzed
An AI software factory is five stages with a gate at each one. The agent is the cheap part. StageWhat it decidesPublished exampleIntakeWhich work is worth startingSentry's Seer scores every incoming issue for actionability firstIsolationWhere the agent runs without collidingStripe boots pre-warmed devboxes in about 10 secondsToolsWhat the agent can reachStripe's Toolshed exposes roughly 500 internal tools over MCPVerificationWhether the change is rightSpotify's LLM judge vetoes about 25% of agent sessionsMerge gateWho is accountableFaire requires two human reviews on agent-authored PRs The short version: every company that made this work built the gates before the fleet. Spotify's Fleetshift shipped in 2023, two years before it had an agent to put in it. Generation scales with spend, review does not, and that asymmetry is the whole design problem. Where Firecrawl fits. Agents need live web context the repo does not carry. Search + scrape for the open web and a curated developer index for code, both behind one MCP block, with prompt injection detection on every fetch. On January 6, 2026, Stephen Toub opened nine pull requests from his phone at 35,000 feet. Seven of them merged. He works on dotnet/runtime, and he wrote up what the experience told him: AI changes the economics of code production. One person with good judgment and a phone can generate PRs faster than a team can review them. That sentence is the entire subject. A single engineer with a coding agent can now saturate a team's review capacity from an airplane seat. The interesting question stopped being how to make agents write code and became how to absorb the output. An AI software factory is the answer companies have converged on, and it is what turns autonomous coding agents from a demo into throughput a team can absorb. This guide breaks it into five stages, each with the published architecture behind it and the configuration to build it. For the broader picture of how AI agents reason, call tools, and pull in web context, start with that primer, then come back here for how a coding-agent fleet actually gets shipped. Assume you have already picked an agent, which we covered in our roundup of the best AI coding agents. The software factory is everything around it. An AI software factory is the system around a coding agent rather than the agent itself. Work arrives from a queue, agents run in isolated workspaces, verification happens automatically, and a human sits at an explicit merge gate. Also called an agentic software factory. The useful distinction is between an agent and a software factory. Running a coding agent on your laptop is an agent: you choose the task, you watch it work, you read the diff, you merge. Everything except the typing is still you, and your attention is the limit. A software factory moves those steps into infrastructure. Nobody decides which issue an agent picks up, because intake rules do. Nobody sets up a workspace, because isolation is provisioned. Nobody checks whether the change compiles, because verification runs before a human is involved at all. The person shows up at the end, on the decisions that carry accountability. Addy Osmani puts it more compactly: A software factory is harnessing loops at scale. The loops he means are the ones we covered in our guide to loop engineering: an agent that runs, checks its own work, and runs again until a verifier says stop. The software factory is the machinery that runs many of them at once without anyone watching. Three properties separate a real software factory from a pile of scripts: It is queue-driven, not prompt-driven. Work enters from issues, alerts, or a Slack channel, and the system decides what is worth starting. Nobody is typing prompts. Environments are disposable. Every agent gets a clean workspace it can destroy, so a bad run costs nothing and parallel runs cannot corrupt each other. Verification runs before review. By the time a diff reaches a person, it has already compiled, passed tests, and been checked for scope. Miss the third and you have not built a software factory. You have built a machine that generates review work faster than you can absorb it, which is the failure mode the rest of this article is organized around avoiding. Guillermo Rauch, Vercel's CEO, made the strategic case when Vercel open sourced its own reference platform for cloud coding agents, and he named the same systems this article draws on: You've heard that companies like Stripe (Minions), Ramp (Inspect), Spotify (Honk), Block (Goose), and others are building their own "AI software factories".
Why? [...] On a business level, the moat of software companies will shift from 'the code they wrote', to the 'means of production' of that code. The alpha is in your factory.