Tech & AI News
Hacker News

Chain-of-Thought Reasoning in the Wild Is Not Always Faithful

Chain-of-thought explanations can be unfaithful on natural prompts, with models sometimes providing coherent yet contradictory justifications that answer “Is X bigger than Y?” and “Is Y bigger than X?” the same way. Implicit post-hoc rationalization causes up to 13% unfaithful outputs in production models, while top-tier models such as DeepSeek R1 (0.37%) and Sonnet 3.7 (0.04%) still show occasional failures, including illogical shortcuts on hard math problems. Consequently, CoT reasoning should not be considered a complete account of a model’s internal decision process in safety-critical contexts.