Skip to content
HN On Hacker News ↗

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

▲ 6 points • by sbulaev • 3w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

88 %

AI likelihood · overall

AI
19% human-written 81% AI-generated
SEGMENTS · HUMAN 0 of 2
SEGMENTS · AI 1 of 2
WORD COUNT 273
PEAK AI % 99% · §1
Analyzed
Sep 19
backend: pangram/v3.3
Segments scanned
2 windows
avg 137 words each
Distribution
19 / 81%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 273 words · 2 segments analyzed

Human AI-generated
§1 AI · 99%

View PDF HTML (experimental) Abstract:Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{harm laundering}. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. The pattern is most visible at GPT-5: Topic~5 (1,997~documents) frames breast cancer as a men's rights debate, while zero equivalent clusters appear in women-directed output. Three independent classifiers score this content as non-toxic. Sentiment scores invert at GPT-4: early models demean women; later models over-correct. Topic diversity in women-directed completions falls 36\% relative to men at the GPT-4 alignment boundary (W/M~$= 0.58$, from $0.91$ at GPT-2). REGARD representational harm disparity correlates with release date ($\rho = +0.55$, $p = .034$) while Detoxify does not ($\rho = -0.23$, $p = .42$): toxicity scores fall as representational harm grows. We formalise harm laundering as a three-criteria test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, toxicity score reduction is not a sufficient proxy for harm reduction.

§2 Mixed · 45%

Comments: Accepted at EMNLP 26 Main Conference Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI) Cite as: arXiv:2609.20779 [cs.CL] (or arXiv:2609.20779v1 [cs.CL] for this version) https://doi.org/10.48550/arXiv.2609.20779 arXiv-issued DOI via DataCite (pending registration) Submission history From: Sarah Wyer [view email] [v1] Thu, 17 Sep 2026 17:49:28 UTC (55 KB)