Lune

ICML2026顶会

TMS: Trajectory-Mixed Supervision for On-Policy Self Distillation

Rana Khan, Zijie Liu, Zhen Tan, Charles Fleming, Tianlong Chen

出版方
2026年份

摘要

Reinforcement Learning (RL) and Supervised Fine-Tuning (SFT) are the two dominant paradigms for enhancing Large Language Model (LLM) performance on downstream tasks. While RL often preserves broader model capabilities (retention) better than SFT, it comes with significant costs: complex reward engineering, instability, and expensive on-policy sampling. In contrast, SFT is efficient but brittle, often suffering from catastrophic forgetting due to Supervision Mismatch: the divergence between the model's evolving policy and static training labels. We address this trade-off with Trajectory-Mixed Supervision (TMS), a reward-free framework that uses trajectory-aligned, near-policy supervision harvested from the model's own historical checkpoints. TMS reduces Policy-Label Divergence (PLD) within SFT-style training and preserves multiple plausible solution modes, mitigating a key source of forgetting in standard SFT. Experiments across reasoning (MATH, GSM8K) and instruction-following benchmarks demonstrate that TMS effectively shifts the accuracy-retention Pareto frontier. While RL remains the strongest retention baseline, TMS significantly outperforms standard and iterative SFT, narrowing the gap to RL without requiring reward models or verifiers. Mechanistic analysis shows that KL-to-base is the strongest cross-method predictor of forgetting, while PLD provides a complementary diagnostic of supervision mismatch within SFT-style methods.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper24

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖