Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
MixedArticle text · 264 words · 4 segments analyzed
View PDF HTML (experimental) Abstract:Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-level score neither characterizes a text nor identifies which AI model wrote it.
We ask whether AI-generated text can be identified one level deeper, from structural signatures: how information is presented, in what order, with what evidence, and in what voice. We replicate StoryScope (Russell et al., 2026), which showed such patterns for AI-generated fiction, on commercial content: 2,250 pre-ChatGPT human blog posts from 268 company domains against 11,250 AI mirrors from five frontier models.
A 214-feature instrument, applied by an LLM and validated in a human gold-annotation session (human-human kappa 0.928, human-model 0.946), detects AI posts from its 187 structural features alone at 98.0 macro-F1 on held-out companies, unchanged (98.1) when every AI post is reworded by its own model. The signal characterizes and attributes: AI posts share a tidy, self-announcing shape, 79.3% are attributed to the correct source against a 16.7% chance rate, and human posts occupy rare structural configurations. All effects replicate StoryScope's, consistent in direction and larger in magnitude.
We release pipeline, instrument, prompts, code, and aggregate artifacts. Comments: 20 pages, 5 figures. Verification artifacts and code: this https URL. v2: corrected description of brief construction and several reported counts; added AI disclosure Subjects: Computation and Language (cs.CL) Cite as: arXiv:2609.15369 [cs.CL] (or arXiv:2609.15369v2 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2609.15369 arXiv-issued DOI via DataCite Submission history From: Jochen Madler [view email] [v1] Mon, 14 Sep 2026 10:55:30 UTC (357 KB) [v2] Thu, 17 Sep 2026 06:51:51 UTC (357 KB)