LLM×MapReduce: Simplified Long-Sequence Processing using Large Language Models
Zihan Zhou, Chong Li, Xinyi Chen, Shuo Wang, Yu Chao, Zhili Li, Haoyu Wang, Qi Shi, Zhixing Tan, Xu Han, Xiaodong Shi, Zhiyuan Liu, Maosong Sun
Abstract
Enlarging the context window of large language models (LLMs) has become a crucial research area, particularly for applications involving extremely long texts. In this work, we propose a novel training-free framework for processing long texts, utilizing a divide-and-conquer strategy to achieve comprehensive document understanding. The proposed LLMMapReduce framework splits the entire document into several chunks for LLMs to read and then aggregates the intermediate answers to produce the final output. The main challenge for divide-and-conquer long text processing frameworks lies in the risk of losing essential long-range information when splitting the document, which can lead the model to produce incomplete or incorrect answers based on the segmented texts. Disrupted long-range information can be classified into two categories: inter-chunk dependency and inter-chunk conflict. We design a structured information protocol to better cope with inter-chunk dependency and an in-context confidence calibration mechanism to resolve inter-chunk conflicts. Experimental results demonstrate that LLMMapReduce can outperform representative open-source and commercial long-context LLMs, and is applicable to several different models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6cf085a7-aff8-408a-a9c7-2f8987df0074Cited by top-tier papers5
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo et al.ACL 2025 · 334 citations
- LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context RewardsGuanzheng Chen, Michael Qizhe Shieh, Lidong BingICLR 2026 · 18 citations
- Benefits and Limitations of Communication in Multi-Agent ReasoningMichael Rizvi-Martel, Satwik Bhattamishra, Neil Rathi, Guillaume Rabusseau et al.ICLR 2026 · 9 citations
- Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent CollaborationZijun Liu, Zhennan Wan, Peng Li, Ming Yan et al.ACL 2026 · 2 citations
- ToM: Leveraging Tree-oriented MapReduce for Long-Context Reasoning in Large Language ModelsJiani Guo, Zuchao Li, Jie Wu, Qianren Wang et al.EMNLP 2025
Builds on3
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsYukang Chen, Shengju Qian, Haotian Tang, Xin Lai et al.ICLR 2024 · 254 citations
- WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-InstructHaipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao et al.ICLR 2025
Related papers
- Hierarchical Context Merging: Better Long Context Understanding for Pre-trained LLMsWoomin Song, Seunghyuk Oh, Sangwoo Mo, Jaehyung Kim et al.ICLR 2024 · 39 citations
- InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context MemoryChaojun Xiao, Pengle Zhang, Xu Han, Guangxuan Xiao et al.NeurIPS 2024 · 223 citations
- When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition FrameworkZach Xu, Shang Zhu, Jue Wang, Junlin Wang et al.ICLR 2026 · 9 citations
- Training-Free Long-Context Scaling of Large Language ModelsChenxin An, Fei Huang, Jun Zhang, Shansan Gong et al.ICML 2024 · 68 citations
- FocusLLM: Precise Understanding of Long Context by Dynamic CondensingZhenyu Li, Yike Zhang, Tengyu Pan, Yutao Sun et al.ACL 2025 · 13 citations
