What are the Essential Factors in Crafting Effective Long Context Multi-Hop Instruction Datasets? Insights and Best Practices
Zhi Chen, Qiguang Chen, Libo Qin, Qipeng Guo, Haijun Lv, Yicheng Zou, Hang Yan, Kai Chen, Dahua Lin
Abstract
Recent advancements in large language models (LLMs) with extended context windows have significantly improved various tasks. To improve long-context capabilities, much work focuses on augmenting LLM's capabilities with synthetic data. Existing methods often leverage the Self-Instruct framework to generate longcontext instruction-tuning data. However, our preliminary experiments show that fewer than 35% of samples generated by Qwen-2 72B are multi-hop, and over 40% exhibit poor quality, limiting comprehensive understanding and further research. To address this, we propose the Multi-agent Interactive Multi-hop Generation (MIMG) framework, which integrates a quality verification agent, a single-hop question generation agent, a multiple question sampling strategy, and a multi-hop question merger agent. This framework significantly improves data quality, with high-quality, multi-hop, and diverse data. Furthermore, we conduct a thorough analysis of document selection, question merging, and validation techniques through extensive experiments across various models. Our results demonstrate that synthetic highquality long-context instruction data can enhance model performance, surpassing even models trained on larger amounts of humanannotated data. Our code and relevant data are available at: https://github.com/ WowCZ/LongMIT . Document 1 Question 11 Answer 11 … Question Semantic Relevance Matrix Multi-hop Question 1 Multi-hop Answer 1 Intra-Document Multi-hop Data Multi-hop Question 2 Multi-hop Answer 2 Inter-Document Multi-hop Data Doc 1 Doc 2 Doc 3 Q11 Q12 Q13 Q14 Q21 Q22 Q23 Q 31 Q 32 Doc 1 Doc 2 Doc 3 Q11 Q12 Q13 Q14
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language ModelsZiyi Yang, Weizhou Shen, Chenliang Li, Ruijun Chen et al.ICLR 2026 · 27 citations
- Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference OptimizationShaohua Duan, Pengcheng Huang, Xinze Li, Zhenghao Liu et al.ACL 2026 · 7 citations
- Revisiting Long-context Modeling from Context Denoising PerspectiveZecheng Tang, Baibei Ji, Juntao Li, Lijun Wu et al.ICLR 2026 · 5 citations
- HopWeaver: Cross-Document Synthesis of High-Quality and Authentic Multi-Hop QuestionsZhiyu Shen, Jiyuan Liu, Yunhe Pang, Yanghui Rao et al.ACL 2026 · 2 citations
- LOGO - Long cOntext aliGnment via efficient preference OptimizationZecheng Tang, Zechen Sun, Juntao Li, Qiaoming Zhu et al.ICML 2025
Builds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
Related papers
- CGMIS: Concept-Graph Based Multi-Hop Instructions Synthesis for Enhancing Long-Context ReasoningZechen Sun, Zecheng Tang, Juntao Li, Wenpeng Hu et al.AAAI 2026
- LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context InstructionsChaochen Gao, Xing Wu, Zijia Lin, Debing Zhang et al.NeurIPS 2025 · 8 citations
- VIGC: Visual Instruction Generation and CorrectionBin Wang, Fan Wu, Xiao Han, Jiahui Peng et al.AAAI 2024 · 95 citations
- MAIN: Mutual Alignment Is Necessary for instruction tuningFanyi Yang, Jianfeng Liu, Xin Zhang, Haoyu Liu et al.EMNLP 2025
- Star-Agents: Automatic Data Optimization with LLM Agents for Instruction TuningHang Zhou, Yehui Tang, Haochen Qin, Yujie Yang et al.NeurIPS 2024 · 21 citations
