XR: Cross-Modal Agents for Composed Image Retrieval
Zhongyu Yang, Wei Pang, Yingfang Yuan
Abstract
Retrieval is being redefined by agentic AI, demanding multimodal reasoning beyond conventional similarity-based paradigms. Composed Image Retrieval (CIR) exemplifies this shift as each query combines a reference image with textual modifications, requiring compositional understanding across modalities. While embedding-based CIR methods have achieved progress, they remain narrow in perspective, capturing limited cross-modal cues and lacking semantic reasoning. To address these limitations, we introduce XR, a training-free multi-agent framework that reframes retrieval as a progressively coordinated reasoning process. It orchestrates three specialized types of agents: imagination agents synthesize target representations through cross-modal generation, similarity agents perform coarse filtering via hybrid matching, and question agents verify factual consistency through targeted reasoning for fine filtering. Through progressive multi-agent coordination, XR iteratively refines retrieval to meet both semantic and visual query constraints, achieving up to a 38% gain over strong training-free and training-based baselines on FashionIQ, CIRR, and CIRCO, while ablations show each agent is essential. Code is available: https://01yzzyu.github.io/xr.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58f24c55-ce57-426a-a842-89aaea693020Cited by top-tier papers2
- SVAgent: Storyline-guided Long Video Understanding via Cross-Modal Multi-Agent CollaborationZhongyu Yang, Zuhao Yang, Shuo Zhan, Tan Yue et al.CVPR 2026 · 5 citations
- Do Vision and Text Cues Exhibit Evidential Coupling? UFO: A Benchmark for Compositional Multimodal Reasoning in Unified ModelsZhongyu Yang, Dannong Xu, Yonghan Zhang, Kefan Chen et al.ICML 2026
Builds on29
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement LearningLakshya A. Agrawal, Shangyin Tan, Dilara Soylu, Noah Ziems et al.ICLR 2026 · 466 citations
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language ModelsZheyuan Liu, Cristian Rodriguez Opazo, Damien Teney, Stephen GouldICCV 2021 · 344 citations
Related papers
- Generative Thinking, Corrective Action: User-Friendly Composed Image Retrieval via Automatic Multi-Agent CollaborationZhangtao Cheng, Yuhao Ma, Jian Lang, Kunpeng Zhang et al.KDD 2025 · 2 citations
- ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image RetrievalTianyu Yang, ChenWei He, Xiangzhao Hao, Tianyue Wang et al.CVPR 2026 · 3 citations
- SDR-CIR: Semantic Debias Retrieval Framework for Training-Free Zero-Shot Composed Image RetrievalYi Sun, Jinyu Xu, Qing Xie, Jiachen Li et al.WWW 2026 · 1 citation
- CaLa: Complementary Association Learning for Augmenting Comoposed Image RetrievalXintong Jiang, Yaxiong Wang, Mengjian Li, Yujiao Wu et al.SIGIR 2024 · 13 citations
- ConText-CIR: Learning from Concepts in Text for Composed Image RetrievalEric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou et al.CVPR 2025
