CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency
Zhanming Shen, Hao Chen, Yulei Tang, Shaolin Zhu, Wentao Ye, Xiaomeng Hu, Haobo Wang, Gang Chen, Junbo Zhao
摘要
Instruction tuning is vital for aligning large language models (LLMs) with human intent, but current methods typically rely on costly human-annotated seed data or powerful external teacher models. While instruction backtranslation techniques reduce this dependency, they remain fundamentally tethered to an initial seed set, which limits full automation, introduces biases, and can lead to inefficient use of unlabeled corpora. In this paper, we propose CYCLE-INSTRUCT, a novel framework that achieves fully seed-free instruction tuning. Inspired by cycle consistency, CYCLE-INSTRUCT employs a dual self-training loop where two models-an answer generator and a question generator-are bootstrapped solely from raw, unlabeled text. These models mutually supervise each other by reconstructing original text segments from their counterpart's generated pseudo-labels, effectively learning from the intrinsic structure of the data without any human-provided seeds. We demonstrate CYCLE-INSTRUCT's efficacy across four diverse data tracks, including general instructionfollowing, domain-specific tasks, dialogue logs, and plain text. Our extensive experiments show that CYCLE-INSTRUCT not only outperforms seed-driven back-translation baselines but also achieves performance comparable to strongly supervised methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Training-Trajectory-Aware Token SelectionZhanming Shen, Jiaqi Hu, Zeyu Qin, Hao Chen 等ICML 2026 · 被引用 1 次
- FedSDR: Federated Self-Distillation with RectificationZiheng Ren, Zhanming Shen, Hao Wang, Ning Liu 等ICML 2026
它引用的顶会 Paper9
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 被引用 887 次
相关 Paper
- Self-Alignment with Instruction BacktranslationXian Li, Ping Yu, Chunting Zhou, Timo Schick 等ICLR 2024 · 被引用 174 次
- Faithful Low-Resource Data-to-Text Generation through Cycle TrainingZhuoer Wang, Marcus D. Collins, Nikhita Vedula, Simone Filice 等ACL 2023 · 被引用 3 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- SelfCodeAlign: Self-Alignment for Code GenerationYuxiang Wei, Federico Cassano, Jiawei Liu, Yifeng Ding 等NeurIPS 2024 · 被引用 79 次
- MAIN: Mutual Alignment Is Necessary for instruction tuningFanyi Yang, Jianfeng Liu, Xin Zhang, Haoyu Liu 等EMNLP 2025
