Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generation
Jiayu Yao, Shenghua Liu, Yiwei Wang, Lingrui Mei, Baolong Bi, Yuyao Ge, Zhecheng Li, Xueqi Cheng
Abstract
Multimodal Retrieval-Augmented Generation (RAG) systems have become essential in knowledge-intensive and open-domain tasks.As retrieval complexity increases, ensuring the robustness of these systems is critical.However, current RAG models are highly sensitive to the order in which evidence is presented, often resulting in unstable performance and biased reasoning, particularly as the number of retrieved items or modality diversity grows.This raises a central question: How does the position of retrieved evidence affect multimodal RAG performance?To answer this, we present the first comprehensive study of position bias in multimodal RAG systems.Through controlled experiments across text-only, imageonly, and mixed-modality tasks, we observe a consistent U-shaped accuracy curve with respect to evidence position.To quantify this bias, we introduce the Position Sensitivity Index (P SI p ) and develop a visualization framework to trace attention allocation patterns across decoder layers.Our results reveal that multimodal interactions intensify position bias compared to unimodal settings, and that this bias increases logarithmically with retrieval range.These findings offer both theoretical and empirical foundations for position-aware analysis in RAG, highlighting the need for evidence reordering or debiasing strategies to build more reliable and equitable generation systems.Our code and experimental resources are available at https://github.com/Theodyy/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e94e1664-91d7-43d1-a947-bf833a46eba0Cited by top-tier papers2
- RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document UnderstandingYinglu Li, Zhiying Lu, Zhihang Liu, Yiwei Sun et al.AAAI 2026 · 2 citations
- Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage OptimizationXu Chu, Guanyu Wang, Zhijie Tan, Xinrong Chen et al.ACL 2026
Builds on9
- RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language ModelsPeng Xia, Kangyu Zhu, Haoran Li, Hongtu Zhu et al.EMNLP 2024 · 39 citations
- DocPrompting: Generating Code by Retrieving the DocsShuyan Zhou, Uri Alon, Frank F. Xu, Zhengbao Jiang et al.ICLR 2023 · 18 citations
- Landmark Embedding: A Chunking-Free Embedding Method For Retrieval Augmented Long-Context Large Language ModelsKun Luo, Zheng Liu, Shitao Xiao, Tong Zhou et al.ACL 2024 · 10 citations
- UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and GenerationXiangyu Zhao, Yuehan Zhang, Wenlong Zhang, Xiao-Ming WuEMNLP 2024 · 7 citations
- MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language ModelsPeng Xia, Kangyu Zhu, Haoran Li, Tianze Wang et al.ICLR 2025 · 5 citations
Related papers
- Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory GraphQiuchen Wang, Shihang Wang, Yu Zeng, Qiang Zhang et al.ICML 2026 · 2 citations
- Positional Bias in Multimodal Embedding Models: Do They Favor the Beginning, the Middle, or the End?Kebin Wu, Fatima AlbreikiAAAI 2026
- HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question AnsweringJoongmin Shin, Gyuho Shim, Jeongbae Park, Jaehyung Seo et al.ACL 2026
- Diagnosing Evidence Utilization in Multimodal Document Question AnsweringDebolena Basak, Digbalay Bose, Koustava Goswami, Maunendra Sankar DesarkarKDD 2026
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document UnderstandingSensen Gao, Shanshan Zhao, Xu Jiang, Lunhao Duan et al.ACL 2026 · 7 citations
