Skip to content
HN On Hacker News ↗

Write Like It's 1866: LLMs Relearn Telegraphese

▲ 94 points • 58 comments • by Theory42 • 2d ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

73 %

AI likelihood · overall

Mixed
22% human-written 78% AI-generated
SEGMENTS · HUMAN 0 of 5
SEGMENTS · AI 3 of 5
WORD COUNT 1,283
PEAK AI % 93% · §1
Analyzed
Oct 7
backend: pangram/v3.3
Segments scanned
5 windows
avg 257 words each
Distribution
22 / 78%
human / AI fraction
Verdict
Mixed
Pangram v3.3

Article text · 1,283 words · 5 segments analyzed

Human AI-generated
§1 AI · 93%

Everything below is reproducible — harness, frozen ledgers, one-click notebook: github.com/Travis42/telegraph-test.Numbers from a 50-passage, ~1,300-question benchmark:Cross-family matrix (readers = foreign models answering from GLM-5.3-Flash’s records; writers = GLM-5.3-Flash answering from theirs)1Savings shown use the lowercase instruction, on each provider’s own meter. Without that word, models write cablese in ALL CAPS and the styling costs 14–19 points: gemma 25.0%, qwen 29.8%, GLM 33.9%. *gpt-5-mini cannot disable reasoning, and compression makes it think — its writes bill about double plain. The one family where this technique does not pay. :ModelRoleplaintext acccablese accrecovery ratio (1.00 = plaintext control)token savingsgemma-4-31breader70.9%77.5%1.09—qwen3.8-27breader72.3%79.4%1.10—nemotron-3-120breader76.1%76.7%1.01—gemma-4-26breader71.8%76.9%1.07—gemma-4-31bwriter——1.0940.4%qwen3.8-27bwriter——1.1048.9%gpt-5-miniwriter——0.9917.7%*GLM-5.3-Flash itself: 48.4% savings with the lowercase instruction, in-family recovery 1.09.2Readers answer 694 anchored questions per condition; writers’ records are read by GLM-5.3-Flash — every GLM call in every run in this post is glm-5.3-flash, the fast variant on the z.ai coding plan. No comparison in the matrix favors plaintext; every ratio sits at 0.99–1.10.The condition ladder — same questions, one variable at a time:ConditionWhat the answering model gets / doesAnchored accuracyvs plaintext controlpassage, answers in plaintextthe full passage91.5%1.00 (reference)passage, answers in cablesefull passage, terse register74.4%0.81└ + transcribed to plaintextthose terse answers, expanded78.8%0.86plaintext recorda plaintext summary75.8%1.00 (reference)cablese recordthe compressed summary82.5%1.09decoded recordcablese record, expanded back81.6%1.08Downstream models, including four families never shown an example, answer questions from the compressed records as well as or better than from plaintext — recovery ratios 0.99–1.10, where 1.00 means “exactly as good as plaintext.” The register is not a construct we invented (LLMs were handed compressed records cold and read them at parity); the capability was already in the weights, inherited from a century and a half of people writing under metered bandwidth.The retention result, interpreted. Cablese records read slightly better than regular responses (the ladder’s 1.09), most likely because compression suppresses copy-the-record phrasing that strict grading penalizes.Every model tested can do this. It’s in their training data — but compression varies from 25% to 49% under identical instructions. That per-model spread is worth knowing if you want to use this technique, and it is knowable.There is a slight information reconstruction cost, but only when transcribing individual answers for users. If the consumer is a model — or the record is simply re-read — the readback ratios are the ones that apply, and they’re all ≥0.99.3Statistically: no comparison favors plaintext; in-family read-back p<0.0001, decoded-record p=0.03.The storage loop’s economics. Expansion gives back about a third of the savings, so a human-readable expanded archive still costs well under plaintext.Notice what this adds up to:No new hardware, no training, no API change, one sentence of instruction. Either your agents get nearly twice the context at the compressed end, or the same work for half the output bill. It is a large optimization that is verified, and as far as I can tell, just not happening.When is compression free?It depends on where the compression step sits. Compress while the model is still composing (answers in cablese, 0.81, partly recoverable at 0.86) and you pay a register tax. But if you compress after the content is settled, into a record a machine will read (1.08–1.09), there is nothing to pay.Cablese belongs between “content settled” and “machine consumption”: scratchpads, memory stores, agent-to-agent handoffs. Nowhere a model is still mid-thought.Models whose reasoning you can disable bill the compression clean. gpt-5-mini’s reasoning is mandatory (the API refuses to turn it off), and cablese makes it think roughly 3× harder — its writes cost double despite the shorter text. The lesson is that you must compress with models whose thinking you can control.So:The expensive half of every LLM bill is optional overhead for machine-consumed text. Output tokens cost 3–5× input tokens, and for traffic a model writes for another model, half of it comes off at any major API — one sentence of instruction, no training, no setup (mandatory-reasoning models excepted).Agent memory just got nearly twice as capacious. Store scratchpads and summaries in cablese: models — the intended readers — read them at parity or better, in every family tested. Expand back to plaintext only when a human actually looks, and even then the archive costs less than plain.Model choice matters as much as the instruction. The identical instruction compresses gemma’s records by 40% and Qwen’s by 49% on their own meters — compressibility is a measurable, per-model property, and at scale that spread is real money. The Telegraph Test is a way of determining how good models are at this kind of compression.It probably can’t be walled off. The register lives in the training data of every family tested. A lab that suppresses it in a frontier model just moves the advantage to open models that still carry it.Auditability survives the compression.

§2 Mixed · 32%

Unlike the emergent agent protocols, cablese is human-readable, fixed by convention, and decodable on demand — the tokens shrink without losing the audit trail.The Telegraph era, where every word was meteredIn 1866, sending a message across the new transatlantic cable cost $100 — for ten words4$10 a word, ten-word minimum, roughly $2,600 in today’s money. . That’s real money today; it was serious money then.

§3 AI · 83%

Telegraph companies charged per word, and an entire industry grew out of that price structure.The Great Eastern laying the first successful Atlantic cable (oil painting, National Maritime Museum, public domain). The link that charged $10 a word — and taught a generation of correspondents to write in cablese.Two compression strategies emerged, and they map onto two very different technologies:Codebooks. Publishers sold massive commercial code dictionaries — Bentley’s ABC Telegraphic Code ran to a thousand pages mapping entire business phrases to single code words.OTTER …might mean*steamship arrived, cargo intact, remit balance.* Firms could cut a 40-word negotiation to a 6-word coded exchange. The codebook is a substitution technology: the message exists in full prose, and a dictionary renames chunks of it.Cablese. The operators and correspondents evolved a written format like:"ARRIVE TUESDAY BRING FUNDS STOP CONFIRM WIFE SAILS FRIDAY." This disciplined style saved words without losing meaning, and it was emergent from the cost of the medium.Over time, this format faded as new technologies like the telephone, fax, email, made the per-word cost premium collapse.

§4 Mixed · 32%

Telegraphese kind of survives wherever metering in terms of cost or just the time it takes to tap out a message is still onerous: 160-character SMS begat a whole new abbreviation culture; early Twitter did it again.Well guess what?Tokens are metered words againAn LLM API bill is a telegraph bill.

§5 AI · 77%

You pay per token.So we ran both strategies from ye olden days against modern models:The Codebook: take existing text, substitute code words from a fixed dictionary.Result: about 10% savings. Most prose isn't dictionary-shaped, and the substitution can't remove words the author already wrote — it can only rename them. CableseInstruct the model, “Write a complete record in telegraphese; drop articles and filler; abbreviate; keep every fact, number, and proper noun verbatim — in lowercase, not all caps.”Why does this work across models?The register is already in the weights — an artifact of training data. Telegraph cables, codebooks, and cablese’s cousins across the broader “telegram style” family (headlinese, teletype style, note-taking, SMS abbreviation) are all in the “all of human knowledge” corpora that most models share. Two observations back this. First, models produce fluent, conventionally-shaped cablese from a one-sentence instruction — no examples, no codebook. Second, the readers in the cross-family matrix never saw even that instruction: they were handed compressed records cold and read them at parity, across four model families. Shared zero-shot fluency like that is hard to explain unless the convention is latent in shared human text.5Boring caveat: I have not run the strict control — the same compression instruction stripped of the historical framing (“write as tersely as possible”) — to test whether generic terseness produces equally legible compression.