MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training
Hui Huang, Jiaheng Liu, Yancheng He, Shilong Li, Bing Xu, Conghui Zhu, Muyun Yang, Tiejun Zhao
摘要
Complex instruction-following with elaborate constraints is imperative for Large Language Models (LLMs). While existing methods have constructed data for complex instruction alignment, they all rely on a more advanced model, especially GPT-4, limiting their application. In this paper, we propose a Multi-granularity Self-Contrastive Training (MuSC) framework, to improve the complex instruction alignment without relying on a stronger model. Our method is conducted on both coarse and fine granularity. On coarse-granularity, we construct constraint-aware preference data based on instruction decomposition and recombination. On fine-granularity, we perform tokenaware preference optimization with dynamic token-level supervision. Our method is evaluated on open-sourced models, and experiment results show our method achieves significant improvement on both complex and general instruction-following benchmarks, surpassing previous self-alignment methods 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Incentivizing Reasoning for Advanced Instruction-Following of Large Language ModelsYulei Qin, Gang Li, Zongyi Li, Zihan Xu 等NeurIPS 2025 · 被引用 17 次
- VIGIL: Defending LLM Agents Against Tool-Stream Injection via Verify-Before-CommitJunda Lin, Zhaomeng Zhou, Zhi Zheng, Shuochen Liu 等ACL 2026 · 被引用 7 次
- Instructions are all you need: Self-supervised Reinforcement Learning for Instruction FollowingQingyu Ren, Qianyu He, Powei Chang, Jie Zeng 等ACL 2026 · 被引用 6 次
- Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction FollowingChenyang Wang, Liang Wen, Shousheng Jia, Xiangzheng Zhang 等AAAI 2026 · 被引用 5 次
- AIR: Complex Instruction Generation via Automatic Iterative RefinementWei Liu, Yancheng He, Yu Li, Hui Huang 等EMNLP 2025
它引用的顶会 Paper15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 被引用 1,203 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
相关 Paper
- IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference OptimizationXinghua Zhang, Haiyang Yu, Cheng Fu, Fei Huang 等ACL 2025 · 被引用 29 次
- MAIN: Mutual Alignment Is Necessary for instruction tuningFanyi Yang, Jianfeng Liu, Xin Zhang, Haoyu Liu 等EMNLP 2025
- InstEmb: Instruction-Following Embeddings through Glimpses of the FutureTianhao Gao, Jun Fang, Xiaohui Zhang, Zhiyuan Liu 等ICML 2026
- UltraIF: Advancing Instruction Following from the WildKaikai An, Li Sheng, Ganqu Cui, Shuzheng Si 等EMNLP 2025 · 被引用 1 次
- LIONs: An Empirically Optimized Approach to Align Language ModelsXiao Yu, Qingyang Wu, Yu Li, Zhou YuEMNLP 2024
