Antidistillation Sampling
Yash Savani, Asher Trockman, Zhili Feng, Yixuan Even Xu, Avi Schwarzschild, Alexander Robey, Marc Finzi, Zico Kolter
摘要
Frontier models that generate extended reasoning traces inadvertently produce token sequences that can facilitate model distillation. Recognizing this vulnerability, model owners may seek sampling strategies that limit the effectiveness of distillation without compromising model performance. Antidistillation sampling provides exactly this capability. By strategically modifying a model's next-token probability distribution, antidistillation sampling poisons reasoning traces, rendering them significantly less effective for distillation while preserving the model's utility. Our code is available at https://github.com/locuslab/ antidistillation-sampling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz 等ICML 2023 · 被引用 854 次
相关 Paper
- Protecting Language Models Against Unauthorized Distillation through Trace RewritingXinhang Ma, William Yeoh, Ning Zhang, Yevgeniy VorobeychikACL 2026 · 被引用 6 次
- Antidistillation FingerprintingYixuan Xu, John Kirchenbauer, Yash Savani, Asher Trockman 等ICML 2026 · 被引用 4 次
- ReasMark: A Robust Watermark for Attributing LLM Reasoning Under Knowledge Distillation AttacksPeizhuo Lv, Ruihua Zhou, Yunpeng Li, Ruigang Liang 等ACL 2026
- Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning DistillationKaiyuan Liu, Shaotian Yan, Rui Miao, Bing Wang 等ICLR 2026 · 被引用 7 次
- Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM ReasoningShuyao Xu, Cheng Peng, Jiangxuan Long, Weidi Xu 等ACL 2026 · 被引用 3 次
