Micro-Macro Retrieval: Reducing Long-Form Hallucination in Large Language Models
Yujie Feng, Jian Li, Zhihan Zhou, Pengfei Xu, Yujia Zhang, xiaoyu li, Xiaohui Zhou, Alan Zhao, Xi Chen, Xiao-Ming Wu
Abstract
Large Language Models (LLMs) achieve impressive performance across many tasks but remain prone to hallucination, especially in long-form generation where redundant retrieved contexts and lengthy reasoning chains amplify factual errors. Recent studies highlight a critical phenomenon: the closer key information appears to the model outputs, the higher the factual accuracy. However, existing retrieval-augmented language models (RALMs) lack effective mechanisms to ensure this proximity -external evidence is injected into reasoning via multi-turn retrieval, but this cannot ensure key information stays close to the outputs. We propose Micro-Macro Retrieval (M 2 R), a novel retrieve-while-generate framework to fill this gap. At the macro level, M 2 R retrieves coarse-grained evidence from external sources; at the micro level, it extracts essential results from a key information repository built during reasoning and reuses them while generating answers. This design directly addresses the key-information-to-output proximity bottleneck, effectively reducing hallucination in long-form tasks. M 2 R is trained with a curriculum learning-based reinforcement learning strategy using customized rulebased rewards, enabling stable acquisition of retrieval and grounding skills. Extensive experiments across different benchmarks demonstrate the effectiveness of M 2 R, especially in lengthy-context settings. However, RALMs are far from solving hallucination in long-form generation (Liu et al., 2025b; Chang et al., 2025b). A key challenge, which we refer to as Lost in Lengthy Contexts, arises when key evidence is obscured in long contexts. This challenge manifests in two aspects. First, retrieved results are often lengthy, and the redundant information makes it difficult for the model to capture * Equal contribution. † Corresponding author.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on26
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- RECOMP: Improving Retrieval-Augmented LMs with Context Compression and Selective AugmentationFangyuan Xu, Weijia Shi, Eunsol ChoiICLR 2024 · 260 citations
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 187 citations
- Long-form factuality in large language modelsJerry Wei, Chengrun Yang, Xinying Song, Yifeng Lu et al.NeurIPS 2024 · 182 citations
Related papers
- ARL2: Aligning Retrievers with Black-box Large Language Models via Self-guided Adaptive Relevance LabelingLingxi Zhang, Yue Yu, Kuan Wang, Chao ZhangACL 2024
- LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question AnsweringQingfei Zhao, Ruobing Wang, Yukuo Cen, Daren Zha et al.EMNLP 2024 · 13 citations
- Boosting Retrieval-Augmented Generation with Generation-Augmented Retrieval: A Co-Training ApproachYubao Tang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.SIGIR 2025 · 2 citations
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun et al.EMNLP 2023 · 315 citations
- RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within GenerationXiaoxi Li, Jiajie Jin, Yujia Zhou, Yongkang Wu et al.ACL 2025
