Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Siyan Zhao, Zhihui Xie, Mengchen Liu, Jing Huang, Guan Pang, Feiyu Chen, Aditya Grover
Abstract
Knowledge distillation improves large language model (LLM) reasoning by compressing the knowledge of a teacher LLM to train smaller LLMs. On-policy distillation advances this approach by having the student sample its own trajectories while a teacher LLM provides dense token-level supervision, addressing the distribution mismatch between training and inference in off-policy distillation methods. However, onpolicy distillation typically requires a separate, often larger, teacher LLM and does not explicitly leverage ground-truth solutions available in reasoning datasets. Inspired by the intuition that a sufficiently capable LLM can rationalize external privileged reasoning traces and teach its weaker self, we introduce On-Policy Self-Distillation (OPSD), a learning algorithm where a single LLM acts as both teacher and student with different contexts. The teacher policy conditions on privileged information (e.g., verified reasoning traces) while the student policy sees only the question; training minimizes the pertoken divergence between these distributions over the student's own rollouts. We demonstrate the efficacy of our method on multiple mathematical reasoning benchmarks, achieving superior token efficiency compared to reinforcement learning methods and better performance over offpolicy distillation methods. Code repo: https: //github.com/siyan-zhao/OPSD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1343bccf-bb32-43b7-9908-41cb077807d0Cited by top-tier papers10
- Self-Distillation Enables Continual LearningIdan Shenfeld, Mehul Damani, Jonas Hübotter, Pulkit AgrawalICML 2026 · 159 citations
- Privileged Information Distillation for Language ModelsEmiliano Penaloza, Dheeraj Vattikonda, Nicolas Gontier, Alexandre Lacoste et al.ICML 2026 · 61 citations
- Expanding the Capabilities of Reinforcement Learning via Text FeedbackYuda Song, Lili Chen, Fahim Tajwar, REMI MUNOS et al.ICML 2026 · 41 citations
- Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM TrainingMiaosen Zhang, Yishan Liu, Shuxia Lin, Qi Dai et al.ICML 2026 · 5 citations
- Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level MathGuijin Son, Donghun Yang, Hitesh Patel, Hyunwoo Ko et al.ICML 2026 · 1 citation
Builds on21
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu et al.ICLR 2024 · 637 citations
Related papers
- Entropy-Aware On-Policy Distillation of Language ModelsWoogyeol Jin, Taywon Min, Yongjin Yang, Dennis Wei et al.ICML 2026 · 91 citations
- Reinforcement Learning via Self-DistillationJonas Hübotter, Frederike Lübeck, Lejs Behric, Anton Baumann et al.ICML 2026
- Reinforcement-aware Knowledge Distillation for LLM Reasoningzhaoyang zhang, Shuli Jiang, Yantao Shen, Yuting Zhang et al.ICML 2026
- Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge DistillationMinsang Kim, Seung Jun BaekICLR 2026 · 15 citations
- VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy DistillationWalid Bousselham, Hilde Kuehne, Cordelia SchmidCVPR 2026 · 31 citations
