Skip to content
HN On Hacker News ↗

Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

▲ 65 points 38 comments by florianherrengt 4d ago HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

3 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 320
PEAK AI % 3% · §1
Analyzed
Aug 19
backend: pangram/v3.3
Segments scanned
1 windows
avg 320 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 320 words · 1 segments analyzed

Human AI-generated
§1 Human · 3%

View PDF HTML (experimental) Abstract:Recent studies indicate that when faced with explicit biases in prompts, models often omit mentioning these biases in their Chain-of-Thought (CoT) output, revealing that verbalized reasoning can give an incorrect picture of how models arrive at conclusions (unfaithfulness). In this work, we show that unfaithful CoT also occurs on naturally worded, non-adversarial prompts without adding artificial biases or editing model outputs. We find that when separately presented with the questions "Is X bigger than Y?" and "Is Y bigger than X?", models sometimes produce superficially coherent arguments to justify systematically answering Yes to both or No to both, despite the contradiction. We present preliminary evidence that this is due to models' implicit biases towards Yes or No, labeling this Implicit Post-Hoc Rationalization. Our results reveal rates up to 13% for production models, and while frontier models are more faithful, none are entirely so, including thinking models like DeepSeek R1 (0.37%) and Sonnet 3.7 with thinking (0.04%). We also investigate Unfaithful Illogical Shortcuts, where models use subtly illogical reasoning to make speculative answers to hard math problems seem rigorously proven. Our findings indicate that while CoT can be useful for assessing outputs, it is not a complete account of the internal process that produced the model's answer and should be used with caution in agentic or safety-critical settings. Comments: Published at the 43rd International Conference on Machine Learning (ICML 2026) Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) Cite as: arXiv:2503.08679 [cs.AI] (or arXiv:2503.08679v6 [cs.AI] for this version) https://doi.org/10.48550/arXiv.2503.08679 arXiv-issued DOI via DataCite Submission history From: Iván Arcuschin [view email] [v1] Tue, 11 Mar 2025 17:56:30 UTC (4,311 KB) [v2] Thu, 13 Mar 2025 17:49:58 UTC (4,348 KB) [v3] Wed, 19 Mar 2025 19:20:42 UTC (4,349 KB) [v4] Tue, 17 Jun 2025 17:59:57 UTC (2,337 KB) [v5] Fri, 29 May 2026 17:38:22 UTC (2,378 KB) [v6] Tue, 16 Jun 2026 17:36:22 UTC (2,378 KB)