R-TOFU: Unlearning in Large Reasoning Models
Sangyeon Yoon, Wonje Jeung, Albert No
Abstract
Large Reasoning Models (LRMs) embed private or copyrighted information not only in their final answers but also throughout multistep chain-of-thought (CoT) traces, making reliable unlearning far more demanding than in standard LLMs. We introduce Reasoning-TOFU (R-TOFU), the first benchmark tailored to this setting. R-TOFU augments existing unlearning tasks with realistic CoT annotations and provides step-wise metrics that expose residual knowledge invisible to answer-level checks. Using R-TOFU, we carry out a comprehensive comparison of gradient-based and preference-optimization baselines and show that conventional answer-only objectives leave substantial forget traces in reasoning. We further propose Reasoned IDK, a preferenceoptimization variant that preserves coherent yet inconclusive reasoning, achieving a stronger balance between forgetting efficacy and model utility than earlier refusal styles. Finally, we identify a failure mode: decoding variants such as ZeroThink and LessThink can still reveal forgotten content despite seemingly successful unlearning, emphasizing the need to evaluate models under diverse decoding settings. Together, the benchmark, analysis, and new baseline establish a systematic foundation for studying and improving unlearning in LRMs while preserving their reasoning capabilities. We release R-TOFU and code at https://ai-isl.github.io/r-tofu . LessThink ZeroThink Question Question
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Rethinking Benign Relearning: Syntax as the Hidden Driver of Unlearning FailuresSangyeon Yoon, Hyesoo Hong, Wonje Jeung, Albert NoICLR 2026 · 3 citations
- CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference OptimizationJunyi Li, Yongqiang Chen, Ningning DingACL 2026 · 1 citation
- SEPS: A Separability Measure for Robust Unlearning in LLMsWonje Jeung, Sangyeon Yoon, Albert NoEMNLP 2025
Builds on13
- Large Language Model UnlearningYuanshun Yao, Xiaojun Xu, Yang LiuNeurIPS 2024 · 365 citations
- In-Context Unlearning: Language Models as Few-Shot UnlearnersMartin Pawelczyk, Seth Neel, Himabindu LakkarajuICML 2024 · 217 citations
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee et al.ICLR 2023 · 158 citations
- BadChain: Backdoor Chain-of-Thought Prompting for Large Language ModelsZhen Xiang, Fengqing Jiang, Zidi Xiong, Bhaskar Ramasubramanian et al.ICLR 2024 · 98 citations
- Unlearn What You Want to Forget: Efficient Unlearning for LLMsJiaao Chen, Diyi YangEMNLP 2023 · 40 citations
Related papers
- Towards Reasoning-Preserving Unlearning in Multimodal Large Language ModelsHongji Li, Manjiang Yu, Junchi Yao, PRIYANKA SINGH et al.CVPR 2026 · 3 citations
- STaR: Sensitive Trajectory Regulation for Unlearning in Large Reasoning ModelsJingjing Zhou, Gaoxiang Cong, Li Su, Liang LiAAAI 2026
- Leak@: Unlearning Does Not Make LLMs Forget Under Probabilistic DecodingHadi Reisizadeh, Jiajun Ruan, Yiwei Chen, Soumyadeep Pal et al.ICML 2026
- Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning SkillsChangsheng Wang, Chongyu Fan, Yihua Zhang, Jinghan Jia et al.EMNLP 2025
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning StepsMartin Tutek, Fateme Hashemi Chaleshtori, Ana Marasovic, Yonatan BelinkovEMNLP 2025 · 37 citations
