Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
Shaohua Duan, Pengcheng Huang, Xinze Li, Zhenghao Liu, Xiaoyuan Yi, Yukun Yan, Shuo Wang, Yu Gu, Ge Yu, Maosong Sun
Abstract
Long-context modeling is critical for a wide range of real-world tasks, including longcontext question answering, summarization, and complex reasoning tasks. Recent studies have explored fine-tuning Large Language Models (LLMs) with synthetic data to enhance their long-context capabilities. However, the effectiveness of such approaches is often limited by the low diversity and factual inconsistencies in the generated data. To address these challenges, we propose LongMab, a novel framework that leverages a Multi-Armed Bandit (MAB) rollout strategy to identify the most informative chunks from the given long context for sampling high-quality and diverse responses and constructing preference data pairs for Direct Preference Optimization (DPO) training. Specifically, we treat context chunks as arms of MAB, select chunks based on their expected reward scores to input into LLMs to generate responses, and iteratively update these scores based on reward feedback. Both exploration and exploitation during the rollout process enable the LLM to focus on the most relevant context segments, thereby generating and collecting high-quality and diverse responses. Experimental results on both Llama and Qwen show the effectiveness of LongMab by achieving more than a 4% improvement on long-context reasoning benchmarks. All data and code will be released on https://github. com/NEUIR/LongMab-PO . * indicates equal contribution. † indicates corresponding author. Query: What group of schools is the university where Michael Berland studied a member of? Answer: Five Colleges
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c49398c0-117c-447a-b4ef-89d07b4f8ff8Cited by top-tier papers3
- Autoencoding-Free Context Compression for LLMs via Contextual Semantic AnchorsXin Liu, Runsong Zhao, Pengcheng Huang, Xinyu Liu et al.ICLR 2026 · 16 citations
- Learning to Focus: Causal Attention Distillation via Gradient-Guided Token PruningYiju Guo, Wenkai Yang, Zexu Sun, Ning Ding et al.NeurIPS 2025 · 14 citations
- ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented GenerationPengcheng Huang, Zhenghao Liu, Yukun Yan, Haiyan Zhao et al.NeurIPS 2025 · 11 citations
Builds on18
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsYukang Chen, Shengju Qian, Haotian Tang, Xin Lai et al.ICLR 2024 · 254 citations
Related papers
- CollabLLM: From Passive Responders to Active CollaboratorsShirley Wu, Michel Galley, Baolin Peng, Hao Cheng et al.ICML 2025
- LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context MultitasksYushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng et al.ACL 2025
- Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined DataZhenqing Ling, Daoyuan Chen, Liuyi Yao, Qianli Shen et al.NeurIPS 2025 · 14 citations
- LLM Collaboration with Multi-Agent Reinforcement LearningShuo Liu, Zeyu Liang, Xueguang Lyu, Christopher AmatoAAAI 2026
- M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language ModelsWai-Chung Kwan, Xingshan Zeng, Yufei Wang, Yusen Sun et al.ACL 2024 · 3 citations
