Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
Tommaso Green, Martin Gubri, Haritz Puerto, Sangdoo Yun, Seong Joon Oh
Abstract
We study privacy leakage in the reasoning traces of large reasoning models used as personal agents. Unlike final outputs, reasoning traces are often assumed to be internal and safe. We challenge this assumption by showing that reasoning traces frequently contain sensitive user data, which can be extracted via prompt injections or accidentally leak into outputs. Through probing and agentic evaluations, we demonstrate that test-time compute approaches, particularly increased reasoning steps, amplify such leakage. While increasing the budget of those test-time compute approaches makes models more cautious in their final answers, it also leads them to reason more verbosely and leak more in their own thinking. This reveals a core tension: reasoning improves utility but enlarges the privacy attack surface. We argue that safety efforts must extend to the model's internal thinking, not just its outputs. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54325a63-57db-4488-9c28-614fb63da2efCited by top-tier papers6
- PrivAct: Internalizing Contextual Privacy Preservation via Multi-Agent Preference TrainingYuhan Cheng, Hancheng Ye, Hai Li, Jingwei Sun et al.ICML 2026 · 3 citations
- MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language ModelsXiongtao Sun, HUI LI, Jiaming Zhang, Yujie Yang et al.ICML 2026 · 3 citations
- Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language ModelsAnmol Goel, Cornelius Emde, Seong Joon Oh, Sangdoo Yun et al.ACL 2026 · 1 citation
- CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference OptimizationJunyi Li, Yongqiang Chen, Ningning DingACL 2026 · 1 citation
- MINIM: Privacy-Aware Minimal View for Agents via Trusted Local SanitizationHexuan Yu, Chaoyu Zhang, Heng Jin, Shanghao Shi et al.ICML 2026
Builds on8
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- ProPILE: Probing Privacy Leakage in Large Language ModelsSiwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri et al.NeurIPS 2023 · 229 citations
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity TheoryNiloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov et al.ICLR 2024 · 198 citations
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksJing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur et al.ACL 2024 · 25 citations
- AirGapAgent: Protecting Privacy-Conscious Conversational AgentsEugene Bagdasarian, Ren Yi, Sahra Ghalebikesabi, Peter Kairouz et al.CCS 2024 · 5 citations
Related papers
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha MomentsYuquan Wang, Mi Zhang, Yining Wang, Geng Hong et al.ACL 2026 · 2 citations
- Unveiling Privacy Risks in LLM Agent MemoryBo Wang, Weiyi He, Shenglai Zeng, Zhen Xiang et al.ACL 2025
- Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning SkillsChangsheng Wang, Chongyu Fan, Yihua Zhang, Jinghan Jia et al.EMNLP 2025
- From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 ModelsJue Zhang, Qingwei Lin, Saravan Rajmohan, Dongmei ZhangEMNLP 2025
- Reasoning Traces Shape Outputs but Models Won't Say SoYijie Hao, Lingjie Chen, Ali Emami, Joyce C. HoACL 2026
