Skip to content
HN On Hacker News ↗

Pseudpocalypse

▲ 151 points 86 comments by surprisetalk 1mo ago HN discussion ↗

Pangram verdict · v3.3

We believe that this document is fully human-written

1 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 5 of 5
SEGMENTS · AI 0 of 5
WORD COUNT 1,667
PEAK AI % 0% · §3
Analyzed
Jul 16
backend: pangram/v3.3
Segments scanned
5 windows
avg 333 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 1,667 words · 5 segments analyzed

Human AI-generated
§1 Human · 0%

DYNOMIGHT best topics follow about Here’s a conjecture: If you put any significant amount of text on the internet under different names, those identities can be linked using only the text itself. This is possible (I conject) because of the statistical “fingerprint” you leave in everything you write. Imagine a website where you can paste in some brand-new text someone just wrote. In return, the website provides links to all the text that writer has ever published under any name. It’s not perfect, but it’s pretty good. As far as I know, no such website exists—at least not on the public internet. But I suspect it’s possible and will soon become easy. This will pose some difficulty for pseudonymous blogging. Note: I wrote most of this essay in mid-2025, after which I idiotically sat on it for a year tinkering with theorem statements that none of you will read.1 In the meantime, LLMs have gotten much better at guessing authors from text. (Given the first 1000 words of a draft of this post, Claude 4.8 knows it’s me.) Still, I think we’re just getting started. I expect to see increasingly obscure writers identified from increasingly small bits of text. I expect that this work even when people are writing in a different register or about unrelated subjects. And I expect that everything I’ve ever written under any pseudonym will soon be linked to my genuine-nym.2 A stronger conjecture is that we’re heading towards a sort of generalized pseudpocalypse. Perhaps, in the future, if you interact with the world through essentially any high-bandwidth channel, then you identify yourself. Say you wear a mask in public and only speak by sub-vocalizing into a voice changer. That’s fine, you’ll still be identified using your body shape, gait, or chemical signature. Or say you don’t like your car being tracked everywhere, so you stop carrying a phone and you somehow convince lawmakers to ban license plates. No problem, your car will still be tracked using tiny scratches or unique pinging sounds from the engine. Or say you don’t like being tracked on the internet, so you lock down your browser profile, buy stuff only with Monero, and connect through a chain of three VPNs. That’s OK. You’ll still be identified through how you wiggle your finger as you scroll down the page.

§2 Human · 0%

We’re all just too unique, and the information theoretic limit is coming for us. Starting bits Let’s start from first principles. Imagine that at birth, everyone is assigned a random binary string. Whenever you post anything on the internet, you’re required to sign it with that string. If the strings are very short, like 0110, then lots of other people will have the same one as you. But if the strings are very long, then yours would almost certainly be unique and it would be trivial to link all your pseudonyms. Where’s the transition point? If you only know that the author is currently alive and living somewhere in the Anglosphere, it’s around 29 bits. That’s because if there are K digits, then there are 2ᴷ possible binary strings, and if K = 28.86, then 2ᴷ ≈ 490,000,000 is the number of currently-alive Anglosphere-dwellers. If the strings have fewer than 29 bits, then someone else will probably share your string. If they have more than 29 bits, then your string is probably unique. We don’t (yet?) have to sign the things we write with immutable government-issued strings. But the way you write still provides lots of clues about you by way of your tone, personality, word choice, and so on. Theoretically speaking, I think it has to be possible to link the identities of anyone who writes enough. Imagine again that everyone is assigned a random binary string at birth, but instead of you needing to sign the stuff you write with your string, each time you write a word, there’s some chance that a random bit from your string is revealed and added as a signature to your message. For example, maybe a signature of bit[129]=1 is added, indicating that your string at position 129 has value 1. Think of your string as representing all your writing style quirks, and a bit being revealed as representing when you write something that reveals a preference. For example, maybe bit 18 indicates if you prefer to write your em-dashes with hideous spaces — like this — or without spaces—like this. If you use an em-dash, that bit is revealed. So imagine you’ve written a lot under Pseudonym A, enough that the full bit-string has been revealed.

§3 Human · 0%

Maybe it’s this: Pseudonym A: 110000001111001101110000100001 010100100101011110111001101000 100111110010101001101010111010 Now say you start writing under Pseudonym B. Initially, none of the bits will be known: Pseudonym B: ?????????????????????????????? ?????????????????????????????? ?????????????????????????????? But slowly, you’ll start to leak a few bits: Pseudonym B: ?????00???1????1????????1??0?? ?????01??1?????????11????01??? ???1??????1???10?????????????? And eventually you’ll leak a lot of bits: Pseudonym B: ???0?00?1?11??110??1????10?0?? ?10?001?0101?11???11100?101??? ???11?1??010??10?1?01??011???? Now think about this from the perspective of an “attacker” who wants to know if A and B are the same person. Let’s assume they’ve only seen the above bits, and have no information about anyone else. Then here’s what the attacker knows: A and B have revealed K overlapping bits, which all match. Different people have a 50% chance of matching on any given revealed bit. Non-different people have a 100% chance of matching on any given revealed bit. There are 490,000,000 people. Intuitively, if K was 5, then the fact that all bits match wouldn’t prove much, since with 490 million people, lots of people would match on those bits by chance.

§4 Human · 0%

But if K was 70, it’s extremely unlikely that two different people would share all of them, even with such a gigantic pool to start with. It turns out that if there are N other people with random bits, and you pick K of your bits, the probability that someone exists who matches all of them is 1 - (1-2⁻ᴷ)ᴺ. When N is 490 million, that looks like this: Look at that, 29 appears again. (Isn’t math wonderful?) In general, the transition happens around whatever number of bits K makes 2ᴷ ≈ N, namely K = log₂(N). If you reveal significantly fewer than 29 bits under pseudonym B, then it’s almost guaranteed that there’s someone else out there who matches all of them. But if you reveal significantly more than 29 bits, then there’s almost no chance that anyone else exists who matches all of them. So the attacker essentially knows that A and B are the same person. And I stress again: They know that without needing to see anything from the other 490 million people. Of course, we don’t literally leak bits of immutable feature strings as we write. But you can make the model more realistic, and the same issue persists. If you want to reflect that text only provides noisy information about the writer, then you can add noise to the bits before they’re revealed. If you want to reflect that some writing styles are more common than others, then you can make the distribution over bit strings non-uniform. If you want to reflect that certain quirks are more obvious than others, you can give different bits different probabilities of being revealed. All these make the math more complicated. But they don’t change the basic conclusion: If your writing style contains at least 29 bits of information, and you do enough writing, you’re done. That’s my argument that pseudpocalypse is possible. But I don’t just want to claim that it could happen, eventually. I think it is likely to happen, soon, and that the amount of text you need to reveal isn’t very large. To make that argument, we need to get specific: What features do people have that are reflected in their writing? How many bits of information do those features contain? How accurately can those bits be guessed from written text?

§5 Human · 0%

Note: To avoid this turning into a giant information theory lecture, I’ll mostly use words like “bit” and “information” without being 100% fully precise about what they mean. I’m doing that because I expect that most people reading this aren’t definition-of-bit fetishists, and anyway being hyper-technical would obscure the big picture. If you’re an information theory enthusiast and/or skeptical that I know what I’m doing, I refer you to the Section For Skeptical Information Theory Enthusiasts, below. Until then, use your intuition and have faith. Feature space Say you knew nothing about me other than that I wrote the above words. And say you had to guess my age or religion or occupation. You could guess, right? It wouldn’t be perfect, but you’d do much better than you would without being able to read those words. Thus, somehow, those words contain information about my demographic characteristics. So I tried to make a list of similar things that you could plausibly guess from text at least somewhat better then chance. Here’s what I came up with: Age Education Ethnicity Family status Income Marital status Mental health Native language Occupation Physical health Political leanings Region Religious affiliation Sex In the same spirit, if you only read the above words, could you guess how extroverted or conscientious I am? Again, not perfectly. (When I meet people who read this blog, they usually seem surprised I can survive direct sunlight.) But still, I’m sure you’d do OK. So, again, these words contain information about my personality. What features does personality have? The HEXACO model lists six, namely honesty-humility, emotionality, extraversion, agreeableness, conscientiousness, and openness to experience. I suspect those can all be guessed with reasonable accuracy from a long-enough writing sample. But could you guess more? For each of those six factors, the HEXACO model lists four “facets”. In the abstract, trying to guess 6 × 4 = 24 different personality features from text sounds ludicrous, but just look at them: Honesty-humility Sincerity Fairness Greed avoidance Modesty Emotionality Fearfulness Anxiety Dependence Sentimentality Extraversion Social self-esteem Social boldness Sociability Liveliness Agreeableness Forgivingness