Skip to content
HN On Hacker News ↗

Launching Vespper DOCX MCP: 3× faster, 2× cheaper, more accurate

▲ 38 points • 18 comments • by topaztee • 2w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

13 %

AI likelihood · overall

Mixed
88% human-written 12% AI-generated
SEGMENTS · HUMAN 4 of 9
SEGMENTS · AI 3 of 9
WORD COUNT 1,538
PEAK AI % 84% · §5
Analyzed
Sep 28
backend: pangram/v3.3
Segments scanned
9 windows
avg 171 words each
Distribution
88 / 12%
human / AI fraction
Verdict
Mixed
Pangram v3.3

Article text · 1,538 words · 9 segments analyzed

Human AI-generated
§1 Mixed · 40%

Introduction Word documents are everywhere. In domains such as legal, finance, and healthcare, the Word document is the deliverable. Contracts, regulatory submissions, and audit reports get drafted, redlined, and signed in .docx, and companies usually have libraries of Word templates they work with on a regular basis.

§2 Human · 11%

That work is increasingly shifting to agents. Microsoft Copilot and Claude in Word have made in-document AI mainstream, while a growing number of vertical agents; particularly in legal tech—need to interact with .docx files. However, AI agents are still struggling to do great work in Word documents. We spoke with dozens of software engineers, mostly in the legal tech and health care spaces, who said they spend weeks and even months tuning their harnesses to edit Word documents reliably, and now are forced to maintain very complex in-house solutions. In general, there exist several ways for agents to edit Word docs today, which mainly fall into three categories: Letting agents write code that uses low-level SDKs such as python-docx / aspose / Open XML SDK. Connect MCPs such as SuperDoc, Office CLI, safe-docx or Adeu that provide opinionated tools for the agent. Round-trip the Word document through a lossy projection i.e. convert to Markdown/HTML using something like pandoc/mammoth.js, let the agent edit it and convert it back to .docx. The current solutions work on simple cases, but they fall short when it comes to complex scenarios. Before diving deep into the solutions and their drawbacks, let’s first understand what a .docx file is. Problem A .docx file is essentially a ZIP file with a hierarchy of XML files following the OOXML (Office Open XML) spec. Inside the ZIP, there are the following files: document.xml contains the main text, styles.xml defines reusable styles (kind of like a CSS stylesheet), numbering.xml defines list/numbering behavior, and separate XML files store headers, footers, footnotes, relationships, media, and document metadata. These XML files are quite verbose.

§3 AI · 71%

For example, in document.xml, even a short 4–5-sentence paragraph can turn into thousands of tokens once you add styles, metadata, formatting information, run splitting, and XML boilerplate.

§4 Human · 4%

The text users see in Microsoft Word might be split across many XML nodes and might be persisted in a very verbose manner. Here's an interactive widget that shows how a simple .docx file works behind the scenes: This nature of DOCX editing makes it different and trickier from editing code or HTML. With simple files, the text representation is mostly the thing itself and changes are local. If we take Markdown/simple HTML for example, the structure is at least familiar and styles are local. This problem gets much worse for vertical AI agents. Companies like Harvey have already run into it.

§5 AI · 84%

When they rebuilt their document editing system, the diagnosis they landed on was that they'd been asking one agent to be both a legal assistant and a Word state machine at the same time. An agent like Harvey's already has a hard job. It has to read the counterparty's redlines, apply the firm's playbook, check that a defined term means the same thing in clause 3 as it does in clause 27, and catch that the indemnification cap it just changed contradicts the liability section two pages up. That's the work. Splitting runs and chasing numbering references is not, but it competes for the same context window.

§6 Human · 11%

Status quo Going back to the solutions above, each of them have different trade-offs, but they all land in the same place: the agent spends its context budget on Word mechanics instead of the actual task. Low-level libraries (python-docx, aspose, Open XML SDK) Good: intuitive for the agent - these libraries are in the pre-training data. They provide full expressiveness; usually nothing is off-limits. Bad: slow and expensive. An agent working on a long legal document spends most of its time writing and debugging scripts, and simple things like a hyperlink or a tracked change require the agent to do backflips. MCP servers (SuperDoc, Office CLI, safe-docx, Adeu) Good: fast and cheap compared to writing code. Bad: each one is a new DSL the agent has to learn on the fly, from a large surface of tools and options. And they tend to cover the common 80% - the remaining 20% is where real documents live. Round-tripping (DOCX ↔ Markdown/HTML) Good: the agent doesn't have to think about Word at all. It just edits text. Bad: the conversion is lossy in one direction and can't be undone in the other. Markdown in particular can't express the style relationships a document inherits from its template. Of all these approaches, we believe in the third one - round tripping. We believe that by liberating agents from thinking about Word, they can perform better on their tasks. However, a big problem with round-tripping is - how do you create a lossless conversion? Solution Before explaining how the reconciler works, we want to say why we're so bullish on this shape of solution. The inspiration comes from the Infrastructure as Code world. Before IaC tools became popular, developers had to go to their cloud accounts and "click around". They used to spend hours of navigating the AWS/GCP console, clicking buttons, managing infra by hand. And if you were unlucky enough to have several environments that all had to stay in sync (staging, production, QA, customer envs), or you needed to spin up a new one, that quickly became a nightmare. Then Terraform and Pulumi showed up. Now a developer just changes code, and the tool figures out (or a better word - reconciles) what needs to change in the cloud. No need to spend hours in the console anymore. We thought this idea translates nicely to agents and Word docs. The agent edits something readable that is intuitive to it, and something else figures out what that means for the actual file. But like we said, the tools are not there. Tools like pandoc and mammoth.js lose too much on the way, and they don't really "reconcile" anything to the original file, meaning it's a very lossy approach. So in short: we liked the idea, but the tools were not good enough. We set out to find a better way. First question for us was what representation are we going to use. Markdown was an obvious first candidate, but we dropped it quickly. The reason is because it can't express styles or associate them with elements. Instead, we landed on HTML: HTML is structurally close to OOXML: <w:p> → <p>, <w:hyperlink> → <a>, <w:tbl> → <table> CSS associates styles with specific elements, which is roughly how OOXML styles work too That still left the lossiness. Converters (pandoc, mammoth.js) will turn a DOCX into HTML, but none of them produce HTML that's minimal, clean and high-fidelity at the same time, so we built our own from scratch. So in theory the agent can now just edit our HTML, and all that's left is reconciling it back into the original .docx. This is where things get tricky. OOXML has a huge surface: endless elements, options and styles. If you add a proprietary HTML representation on top of it, you get a very long tail of cases to handle. Instead of building a giant reconciliation engine and chasing every edge case with hand-written code, we decided to train a model to be the reconciler. Namely: show it a lot of HTML ↔ OOXML changes and teach it to predict, given an HTML change, what the OOXML counterpart should be.

§7 AI · 72%

That model sits at the center of the solution. It takes the agent's HTML intent and emits valid OOXML for that block. We then diff the result against the original block, compute tracked changes deterministically, patch it into the file and hand it back to the caller.

§8 Human · 7%

Here's an interactive sequence diagram that shows how our solution works end-to-end: The following section explains how we evaluated our solution against alternatives. Evaluation To benchmark our approach, we compared it against 5 other solutions on 279 DOCX editing tasks from the test set of our internal benchmark, each run on two models: GPT 5.6 Sol and GPT 5.6 Terra, both at medium reasoning: MCP servers Vespper MCP (ours) SuperDoc MCP - v0.18.1 Office CLI (via MCP) - v1.0.145 Adeu MCP - v3.0.2 Skills/Harnesses DOCX skill by Anthropic Plain python-docx harness For the harness we used LangChain's create_agent, which makes agent instantiation and model swapping easy and plays well with the official mcp package. We skipped batteries-included harnesses (LangChain's deepagents, Vercel's eve) because we wanted the most amount of control and the least magic happening behind the scenes (e.g automatic compaction, system prompt modifications, etc). Few more notes about the candidates above: Same system prompt for everyone - No per-solution prompt tuning, including ours. We removed the setup work from all candidates - Almost every solution asks the agent to do some housekeeping before it can edit anything: load a skill, open a file, track a session, save and close when it's done.

§9 Mixed · 43%

We wanted to "neutralize" these parts, so we took it off the agent's plate across the board: the DOCX skill is pre-loaded into context, session IDs and file paths are injected behind the scenes, lifecycle tools like SuperDoc's open/save/close are hidden entirely.