Retrieval-Augmented Perception: High-resolution Image Perception Meets Visual RAG
Wenbin Wang, Yongcheng Jing, Liang Ding, Yingjie Wang, Li Shen, Yong Luo, Bo Du, Dacheng Tao
Abstract
High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs). To drive progress beyond the limits of heuristic methods, this paper advances HR perception capabilities of MLLMs by harnessing cutting-edge long-context techniques such as retrieval-augmented generation (RAG). Towards this end, this paper presents the first study exploring the use of RAG to address HR perception challenges. Specifically, we propose Retrieval-Augmented Perception (RAP), a training-free framework that retrieves and fuses relevant image crops while preserving spatial context using the proposed Spatial-Awareness Layout. To accommodate different tasks, the proposed Retrieved-Exploration Search (RE-Search) dynamically selects the optimal number of crops based on model confidence and retrieval scores. Experimental results on HR benchmarks demonstrate the significant effectiveness of RAP, with LLaVA-v1.5-13B achieving a 43% improvement on V * Bench and 19% on HR-Bench. Code is available at https://github.com/DreamMr/RAP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b74c920-5d6e-4c16-b40f-56f3629a80d2Cited by top-tier papers8
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via SpeculationYuhan Liu, Lianhui Qin, Shenji WanICLR 2026 · 6 citations
- MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image UnderstandingFan Yang, Xingping Dong, Xin Yu, Wenhan Luo et al.CVPR 2026 · 1 citation
- ActiveScope: Actively Seeking and Correcting Perception for MLLMsYajing Wang, Chao Bi, Junshu Sun, Shufan Shen et al.ICML 2026 · 1 citation
- The Perceptual Bandwidth Bottleneck in Vision-Language Models: Active Visual Reasoning via Sequential Experimental DesignAnjie Liu, Ziqin Gong, Yan Song, Yuxiang Chen et al.ICML 2026 · 1 citation
- BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained PerceptionGeng Li, Yuxin PengICML 2026
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
Related papers
- VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality DocumentsShi Yu, Chaoyue Tang, Bokai Xu, Junbo Cui et al.ICLR 2025
- Benchmarking Retrieval-Augmented Generation in Multi-Modal ContextsZhenghao Liu, Xingsheng Zhu, Tianshuo Zhou, Xinyi Zhang et al.ACM MM 2025 · 4 citations
- RetroLM: Retrieval-Augmented KVs for Long-Context ProcessingKun Luo, Zheng Liu, Shitao Xiao, Jiabei Chen et al.AAAI 2026
- ViG-RAG: Video-aware Graph Retrieval-Augmented Generation via Temporal and Semantic Hybrid ReasoningZongsheng Cao, Anran Liu, Yangfan He, Jing Li et al.AAAI 2026 · 1 citation
- AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video UnderstandingZhucun Xue, Jiangning Zhang, Xurong Xie, Yuxuan Cai et al.NeurIPS 2025 · 19 citations
