SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented Generation
Xinping Zhao, Dongfang Li, Yan Zhong, Boren Hu, Yibin Chen, Baotian Hu, Min Zhang
Abstract
Recent studies in Retrieval-Augmented Generation (RAG) have investigated extracting evidence from retrieved passages to reduce computational costs and enhance the final RAG performance, yet it remains challenging. Existing methods heavily rely on heuristic-based augmentation, encountering several issues: (1) Poor generalization due to hand-crafted context filtering; (2) Semantics deficiency due to rulebased context chunking; (3) Skewed length due to sentence-wise filter learning. To address these issues, we propose a model-based evidence extraction learning framework, SEER, optimizing a vanilla model as an evidence extractor with desired properties through selfaligned learning. Extensive experiments show that our method largely improves the final RAG performance, enhances the faithfulness, helpfulness, and conciseness of the extracted evidence, and reduces the evidence length by 9.25 times. The code will be available at https://github.com/HITsz-TMG/SEER .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding ModelXinping Zhao, Xinshuo Hu, Zifei Shan, Shouzheng Huang et al.ICLR 2026 · 47 citations
- Improving Retrieval-Augmented Generation through Multi-Agent Reinforcement LearningYiqun Chen, Lingyong Yan, Weiwei Sun, Xinyu Ma et al.NeurIPS 2025 · 47 citations
- Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question AnsweringHuiyao Chen, Yi Yang, Yinghui Li, Meishan Zhang et al.ACL 2026 · 6 citations
- TSVC: Tripartite Learning with Semantic Variation Consistency for Robust Image-Text RetrievalShuai Lyu, Zijing Tian, Zhonghong Ou, Yifan Zhu et al.AAAI 2025 · 2 citations
- ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image ExplorationHaozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu et al.EMNLP 2025 · 1 citation
Builds on18
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
Related papers
- Embedding-Based Context-Aware RerankerYe Yuan, Amin Shabani, Siqi LiuICLR 2026
- MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question AnsweringHui Wu, Haoquan Zhai, Yuchen Li, Hengyi Cai et al.ACM MM 2025
- Enhancing Retrieval-Augmented Generation via Evidence Tree SearchHao Sun, Hengyi Cai, Yuchen Li, Xuanbo Fan et al.ACL 2025 · 7 citations
- Grounding Language Model with Chunking-Free In-Context RetrievalHongjin Qian, Zheng Liu, Kelong Mao, Yujia Zhou et al.ACL 2024
- EAReranker: Efficient Embedding Adequacy Assessment for Retrieval Augmented GenerationDongyang Zeng, Yaping Liu, Wei Zhang, Shuo Zhang et al.NeurIPS 2025
