Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
Hengyuan Zhang, Shiping Yang, Xiao Liang, Chenming Shang, Yuxuan Jiang, Chaofan Tao, Jing Xiong, Hayden Kwok-Hay So, Ruobing Xie, Angel X. Chang, Ngai Wong
摘要
Training student models on synthetic data generated by strong teacher models is a promising way to distilling the capabilities of teachers. However, recent studies show that stronger models are not always optimal teachers, revealing a mismatch between teacher outputs and student learnability. To address this issue, we propose PerSyn (Personalized data Synthesis), a novel synthesis strategy that operates under a new Route then Generate''paradigm to create data tailored to each student model, enabling it to learn more effectively. Specifically, PerSyn first assigns each prompt to its optimal teacher via a query-level router that jointly considers student learnability and teacher response quality. Each teacher then synthesizes data only for its assigned prompts, making the process more efficient than the conventional Generate then Select''paradigm, where all teachers must generate parallel responses for the entire prompt set before constructing the final dataset. Extensive experiments across different model families and scales demonstrate that PerSyn consistently achieves superior or comparable performance to all baselines in instruct tuning and math reasoning settings. Further analysis verifies the effectiveness of PerSyn and offers extra insights to propel future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Beyond Pass@ 1: Self-Play with Variational Problem Synthesis Sustains RLVRXiao Liang, Zhong-Zhi Li, Yeyun Gong, Yelong Shen 等ICLR 2026 · 被引用 57 次
- Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual GuidanceXinrong Chen, Xu Chu, Yingmin Qiu, Hengyuan Zhang 等CVPR 2026 · 被引用 8 次
- Why Supervised Fine-Tuning Fails to Learn: A Systematic Study of Incomplete Learning in Large Language ModelsChao Xue, Yao Wang, Mengqiao Liu, Di Liang 等ACL 2026 · 被引用 5 次
- MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological OrthogonalityPanqi Yang, Haodong Jing, Jiahao Chao, Tingyan Xiang 等ICML 2026 · 被引用 3 次
- JW-SVD: Bridging the Cross-Modal Mismatch in Post-Training MLLM CompressionRunchao Li, Yao Fu, Mu Sheng, Haotian Yu 等ACL 2026
它引用的顶会 Paper8
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- Beyond Pass@ 1: Self-Play with Variational Problem Synthesis Sustains RLVRXiao Liang, Zhong-Zhi Li, Yeyun Gong, Yelong Shen 等ICLR 2026 · 被引用 57 次
- Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time ScalabilityXiao Liang, Zhong-Zhi Li, Zhenghao Lin, Eric Hanchen Jiang 等ACL 2026 · 被引用 5 次
- Multiple Choice Learning of Low-Rank Adapters for Language ModelingVictor Letzelter, Hugo Malard, Mathieu Fontaine, Gaël Richard 等ICML 2026 · 被引用 1 次
- Task Oriented In-Domain Data AugmentationXiao Liang, Xinyu Hu, Simiao Zuo, Yeyun Gong 等EMNLP 2024 · 被引用 1 次
相关 Paper
- SIPDO: Closed-Loop Prompt Optimization via Synthetic Data FeedbackYaoning Yu, Ye Yu, Peiyan Zhang, Kai Wei 等ICLR 2026 · 被引用 7 次
- Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal SamplingHritik Bansal, Arian Hosseini, Rishabh Agarwal, Vinh Q. Tran 等ICLR 2025 · 被引用 1 次
- Learning from Reasoning Failures via Synthetic Data GenerationGabriela Ben Melech Stan, Estelle Aflalo, Avinash Madasu, Vasudev Lal 等AAAI 2026 · 被引用 2 次
- How to Fine-Tune a Reasoning Model? A Teacher–Student Cooperation Framework to Synthesize Student-Consistent SFT DataZixian Huang, Kaichen Yang, Xu Huang, Feiyang Hao 等ICML 2026 · 被引用 5 次
- Montessori-Instruct: Generate Influential Training Data Tailored for Student LearningXiaochuan Li, Zichun Yu, Chenyan XiongICLR 2025
