Antidistillation Sampling
Yash Savani, Asher Trockman, Zhili Feng, Yixuan Even Xu, Avi Schwarzschild, Alexander Robey, Marc Finzi, Zico Kolter
Abstract
Frontier models that generate extended reasoning traces inadvertently produce token sequences that can facilitate model distillation. Recognizing this vulnerability, model owners may seek sampling strategies that limit the effectiveness of distillation without compromising model performance. Antidistillation sampling provides exactly this capability. By strategically modifying a model's next-token probability distribution, antidistillation sampling poisons reasoning traces, rendering them significantly less effective for distillation while preserving the model's utility. Our code is available at https://github.com/locuslab/ antidistillation-sampling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1844dd1-c77d-4734-be4b-9c6e85c10d2cBuilds on18
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
Related papers
- Protecting Language Models Against Unauthorized Distillation through Trace RewritingXinhang Ma, William Yeoh, Ning Zhang, Yevgeniy VorobeychikACL 2026 · 6 citations
- Antidistillation FingerprintingYixuan Xu, John Kirchenbauer, Yash Savani, Asher Trockman et al.ICML 2026 · 4 citations
- ReasMark: A Robust Watermark for Attributing LLM Reasoning Under Knowledge Distillation AttacksPeizhuo Lv, Ruihua Zhou, Yunpeng Li, Ruigang Liang et al.ACL 2026
- Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning DistillationKaiyuan Liu, Shaotian Yan, Rui Miao, Bing Wang et al.ICLR 2026 · 7 citations
- Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM ReasoningShuyao Xu, Cheng Peng, Jiangxuan Long, Weidi Xu et al.ACL 2026 · 3 citations
