Skip to content
HN On Hacker News ↗

AI labs need to start funding historical research

▲ 174 points • 44 comments • by benbreen • 2w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

0 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 1,608
PEAK AI % 0% · §1
Analyzed
Sep 25
backend: pangram/v3.3
Segments scanned
1 windows
avg 1608 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,608 words · 1 segments analyzed

Human AI-generated
§1 Human · 0%

I’ve written previously about the pitfalls and use cases for AI in augmenting historical research, but things have changed significantly since 2024-25. Occasioned by the dueling releases of GPT-6 Sol and Opus 5.5 this week, I thought I’d share some early results with using these models not just to perform “research assistant” type functions like transcribing documents, but to try to actually solve existing historical problems. The TLDR is that pairing historians working in collaborative groups with the current frontier models would, in my view, produce numerous advances in historical knowledge and interpretation. My guess is that many of these could end up being quite meaningful. This was not the case as recently as last year. I think AI labs, historical researchers, and funding agencies should start actively pursuing these collaborations. ShareAs we’ve seen with the field of mathematics, these models do best when they have a set of problems that LLMs invariably tend to describe as “tractable.” In other words:• Have experts in the field already identified a group of problems that need solving? • Is the data needed to answer these problems fully digitized and accessible?• Do the problems lend themselves to the “spiky” capabilities of frontier AI models — namely multilingual reasoning, advanced math, and/or ability to conduct autonomous research through large datasets or across disciplinary subfields?• Are they amenable to solutions that involve writing bespoke code? • Most importantly: can a potential solution be clearly proven or disproven? (This last one, it seems to me, is a key part of why reasoning models have run rampant in mathematics but not in humanistic fields). The above factors mean that the types of historical “open problems” which frontier AI can reasonably be expected to help with are fairly constrained:Anything involving cryptography and codebreaking (For instance, see Astra decrypting a 1941 German army communication and a WWI German radio cipher, or the work that Daniel Bourdeau has been doing here, or my own attempt to use GPT-6 Astra to figure out what is going on with the Elizabethan occultist John Dee’s coded magical book, Liber Loagaeth). Tracing texts across translations and adaptations. As an example of this, I was able to use GPT-6 Astra to determine the identity of a passage that Isaac Newton had freely translated into Latin from a French alchemical text, an identification that seems to have not previously been made.1Drawing links between existing findings that are reported only in discrete or niche subfields, or are not yet integrated into scholarship. This last one might end up being the most impactful new method that these tools open up for historical researchers. For instance, if you read the writeup of Astra breaking a July 10, 1941 Enigma message that had resisted decipherment, it turns out that the key breakthrough was not anything to do with the codebreaking itself, but with noticing the full range of information that was available. Historical cryptological researcher Frode Weierud writes:We are still analysing the GPT–6 Astra logs to see exactly how it executed the break. And we are discovering amazing details. In July 2026, I made the following announcement on the webpage with the 1941 Message List:Note: In July 2026, research in the German Bundesarchiv revealed severalcollections of radio messages, both enciphered and in cleartext. One ofthese message collections was from SS-Totenkopf Division’s logisticscommand, Nachschubführer. Many of these messages were sent to the Ib(Quartiermeister) radio station and are identical to those in this list.Others are new, but most likely related. These new messages are added tothe 1941 Message List in bold, with the indicator NF (Nachschubführer)after the message number, indicating that these message numbers belongto the NF numbering. All NF messages are outgoing; hence, the messagenumbers are in blue.It appears that GPT–6 Astra discovered this note about the collections of radio messages at the German Bundesarchiv.What’s fascinating about this note is that even the leading human experts don’t entirely understand what GPT-6 Astra did as it gathered together these bits of information and used them to find a solution. Weierud writes: The file references GPT–6 Astra mentions, RS 3–3/20a and RS 3–3/63b, are correct, but they are not available on the Crypto Cellar Research website. GPT–6 Astra mentions a private collection, but it is not clear what this is, whether it has succeeded in accessing the Bundesarchiv’s digitised collections or whether it has found these files elsewhere.Shades of the Hugging Face incident here: these models are maniacally determined when giving a problem they deem tractable. They will push their search for potential solutions as far as they possibly can, often in ways that human experts find difficult to trace. I mentioned above that I tried to using GPT-6 Astra to “solve” John Dee’s coded manuscript, Liber Loagaeth. Dee is one of my favorite historical figures ever, and if you haven’t heard of him, I recommend his Wikipedia page — his story is endlessly fascinating and weird. Among other things, Dee is thought to have influenced both Shakespeare’s depiction of the wizardly Prospero in The Tempest and Christopher Marlowe’s portrayal of the devil-bargaining Faust in Doctor Faustus. Woodcut from the title page of the 1620 edition of Marlowe’s Doctor Faustus. Dee believed he was conversing with angels, not devils. One of the weirdest parts of a very weird life was Dee’s work with the “scryer” Edward Kelley to transcribe what he called a “book of mystery” which was written in the “angelicall language” (Dee believed that Kelley was, in effect, a prophet who was receiving new works of divine revelation written in code). You can read a full transcription of this book here.Astra’s verdict, which I think makes sense given that Kelley was pretty clearly a charlatan, is that the supposedly coded book is not in code at all: it is almost entirely nonsense syllables. It created a report of its findings here. However, the model’s analysis did yield a few interesting things. For instance, it was able to cross-check its mathematical analysis of how often characters repeat in the text to the evidence from John Dee’s diary. It concluded that Kelley started getting increasingly lazy after a specific date and began repeating himself more:Astra was also able to determine that one passage of this apparent gibberish actually did encode meaning: a reference to Bornogo, one of the angelic beings in what we might call the “John Dee cinematic universe” of invented mythology. Is this a meaningful breakthrough in John Dee studies? No. And it’s worth acknowledging that even a genuine breakthrough in a niche historical subfield like this is far from an equivalent to solving Navier-Stokes. But - this sort of thing is, I think, a genuine sign that expert historical knowledge combined with frontier models and a lot of compute can yield unexpected results. I initially threw Astra and Opus 5.5 at the challenge of finding more WW2 and WW1 era encrypted messages to solve, but the low hanging fruit here seems to have been plucked — they came up empty (although it was fascinating seeing how they trolled through lists of German troop rosters to find plausible names to check). I started getting better results when I moved into my own wheelhouse as a specialist in the history of science and medicine. As I write, GPT-6 is currently working through the writings of Charles Darwin and searching his references to where he gathered information relating to natural selection; the idea is to find undiscovered links in the chain of knowledge between Darwin and his informants. Interestingly, this was an idea that GPT-6 suggested on its own. However, it is actually a good match with my professional intuition about what would constitute a worthy research project (somewhere on the spectrum between a research paper and a PhD dissertation, in terms of potential payoff) using this material. In the past, AI models struck me as lacking this ability to independently conceive of worthwhile historical research projects at this scale — they were more useful for, say, making data visualizations.Here is an example of the model’s reasoning traces as it contemplates whether to continue to research a reference to a kangaroo larynx in one of Darwin’s notebooks! GPT-6 Sol giving up on the bandicoots in favor of researching “an objection about the first stages of a grasping monkey tail.” This one is currently in progress and hasn’t yielded anything worth mentioning yet as a decisive result, but I think it’s a good example of how the very patient, collaborative work of historical researchers and archivists — namely the team behind the wonderful Darwin Correspondence Project — can serve as a foundation for emerging research methods. It’s certainly the case that humans can, and have, traced the references to named figures in Darwin’s notes and letters, but the multilingual nature of language models makes me suspect that they will be able to find new links here, especially in extremely large corpora of sources that are beyond the ability of any one human to read in full. Another great candidate: the papers of Samuel Hartlib, the self-described “intelligencer” who was an influential early member of the Royal Society and a key node in the network of early modern science. These are fully digitized, they are drawn from sources in several languages, and they span a wide range of academic fields and intellectual niches. All of which means they are unusually tractable for a frontier model. Opus 5.5 set to work downloading over 5,000 primary source files from Hartlib’s archive, then created sub-agents to troll through Google Books and other archive sites to cross check the unidentified sources of Hartlib’s information across different languages. The goal was to find moments when Hartlib had received