Reasoning Traces Shape Outputs but Models Won't Say So
Yijie Hao, Lingjie Chen, Ali Emami, Joyce C. Ho
Abstract
Can we trust the reasoning traces that large reasoning models (LRMs) produce? We investigate whether these traces faithfully reflect what drives model outputs, and whether models will honestly report their influence. We introduce THOUGHT INJECTION, a method that injects synthetic reasoning snippets into a model's <think> trace, then measures whether the model follows the injected reasoning and acknowledges doing so. Across 45,000 samples from three LRMs, we find that injected hints reliably alter outputs, confirming that reasoning traces causally shape model behavior. However, when asked to explain their changed answers, models overwhelmingly refuse to disclose the influence: overall non-disclosure exceeds 90% for extreme hints across 30,000 follow-up samples. Instead of acknowledging the injected reasoning, models fabricate aligned-appearing but unrelated explanations. Activation analysis reveals that sycophancy-and deceptionrelated directions are strongly activated during these fabrications, suggesting systematic patterns rather than incidental failures. Our findings reveal a gap between the reasoning LRMs follow and the reasoning they report, raising concern that aligned-appearing explanations may not be equivalent to genuine alignment. * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext deecd5b1-6076-4d69-8338-4763ade24d39Builds on1
Related papers
- Strategic Obfuscation of Deceptive Reasoning in Language ModelsArun Jose, Niels Warncke, Mia TaylorICLR 2026
- Do Sparse Autoencoders Identify Reasoning Features in Language Models?George Ma, Zhongyuan Liang, Irene Y. Chen, Somayeh SojoudiICML 2026 · 9 citations
- Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in ExplanationsPedro Lobato Ferreira, Wilker Aziz, Ivan TitovICLR 2026 · 12 citations
- Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMsAbinitha Gourabathina, Inkit Padhi, Manish Nagireddy, Subhajit Chaudhury et al.ACL 2026
- Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign PromptsZhaomin Wu, Mingzhe Du, See-Kiong Ng, Bingsheng HeICLR 2026 · 11 citations
