Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generation
Jiayu Yao, Shenghua Liu, Yiwei Wang, Lingrui Mei, Baolong Bi, Yuyao Ge, Zhecheng Li, Xueqi Cheng
摘要
Multimodal Retrieval-Augmented Generation (RAG) systems have become essential in knowledge-intensive and open-domain tasks.As retrieval complexity increases, ensuring the robustness of these systems is critical.However, current RAG models are highly sensitive to the order in which evidence is presented, often resulting in unstable performance and biased reasoning, particularly as the number of retrieved items or modality diversity grows.This raises a central question: How does the position of retrieved evidence affect multimodal RAG performance?To answer this, we present the first comprehensive study of position bias in multimodal RAG systems.Through controlled experiments across text-only, imageonly, and mixed-modality tasks, we observe a consistent U-shaped accuracy curve with respect to evidence position.To quantify this bias, we introduce the Position Sensitivity Index (P SI p ) and develop a visualization framework to trace attention allocation patterns across decoder layers.Our results reveal that multimodal interactions intensify position bias compared to unimodal settings, and that this bias increases logarithmically with retrieval range.These findings offer both theoretical and empirical foundations for position-aware analysis in RAG, highlighting the need for evidence reordering or debiasing strategies to build more reliable and equitable generation systems.Our code and experimental resources are available at https://github.com/Theodyy/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document UnderstandingYinglu Li, Zhiying Lu, Zhihang Liu, Yiwei Sun 等AAAI 2026 · 被引用 2 次
- Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage OptimizationXu Chu, Guanyu Wang, Zhijie Tan, Xinrong Chen 等ACL 2026
它引用的顶会 Paper9
- RULE: Reliable Multimodal RAG for Factuality in Medical Vision Language ModelsPeng Xia, Kangyu Zhu, Haoran Li, Hongtu Zhu 等EMNLP 2024 · 被引用 39 次
- DocPrompting: Generating Code by Retrieving the DocsShuyan Zhou, Uri Alon, Frank F. Xu, Zhengbao Jiang 等ICLR 2023 · 被引用 18 次
- Landmark Embedding: A Chunking-Free Embedding Method For Retrieval Augmented Long-Context Large Language ModelsKun Luo, Zheng Liu, Shitao Xiao, Tong Zhou 等ACL 2024 · 被引用 10 次
- UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and GenerationXiangyu Zhao, Yuehan Zhang, Wenlong Zhang, Xiao-Ming WuEMNLP 2024 · 被引用 7 次
- MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language ModelsPeng Xia, Kangyu Zhu, Haoran Li, Tianze Wang 等ICLR 2025 · 被引用 5 次
相关 Paper
- Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory GraphQiuchen Wang, Shihang Wang, Yu Zeng, Qiang Zhang 等ICML 2026 · 被引用 2 次
- Positional Bias in Multimodal Embedding Models: Do They Favor the Beginning, the Middle, or the End?Kebin Wu, Fatima AlbreikiAAAI 2026
- HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question AnsweringJoongmin Shin, Gyuho Shim, Jeongbae Park, Jaehyung Seo 等ACL 2026
- Diagnosing Evidence Utilization in Multimodal Document Question AnsweringDebolena Basak, Digbalay Bose, Koustava Goswami, Maunendra Sankar DesarkarKDD 2026
- Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document UnderstandingSensen Gao, Shanshan Zhao, Xu Jiang, Lunhao Duan 等ACL 2026 · 被引用 7 次
