Modeling Uncertainty in Composed Image Retrieval via Probabilistic Embeddings
Haomiao Tang, Jinpeng Wang, Yuang Peng, Guanghao Meng, Ruisheng Luo, Bin Chen, Long Chen, Yaowei Wang, Shutao Xia
摘要
Composed Image Retrieval (CIR) enables users to search for images using multimodal queries that combine text and reference images. While metric learning methods have shown promise, they rely on deterministic point embeddings that fail to capture the inherent uncertainty in the input data, in which user intentions may be imprecisely specified or open to multiple interpretations. We address this challenge by refor-mulating CIR through our proposed Co mposed P robabilistic E mbedding (C O PE) framework, which represents both queries and targets as Gaussian distributions in latent space rather than fixed points. Through careful design of probabilistic distance metrics and hierarchical learning objectives, C O PE explicitly captures uncertainty at both instance and feature levels, enabling more flexible, nuanced, and robust matching that can handle polysemy and ambiguity in search intentions. Extensive experiments across multiple benchmarks demonstrate that C O PE effectively quantifies both quality and semantic uncertainties within Com-posed Image Retrieval, achieving state-of-the-art performance on recall rate. Code: https: //github.com/tanghme0w/ACL25-CoPE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image RetrievalZixu Li, Yupeng Hu, Zhiheng Fu, Zhiwei Chen 等ACL 2026 · 被引用 13 次
- Enhancing Partially Relevant Video Retrieval with Hyperbolic LearningJun Li, Jinpeng Wang, Chaolei Tan, Niu Lian 等ICCV 2025 · 被引用 5 次
- Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video RetrievalJun Li, Xuhang Lou, Jinpeng Wang, Yuting Wang 等CVPR 2026 · 被引用 3 次
- Heterogeneous Uncertainty-Guided Composed Image Retrieval with Fine-Grained Probabilistic LearningHaomiao Tang, Jinpeng Wang, Minyi Zhao, Guanghao Meng 等AAAI 2026 · 被引用 1 次
- Imagine with Layout and Sketch: Enhancing Vision-Language Retrieval with Dual-Stream Multi-Modal Query RefinementGuanghao Meng, Jinpeng Wang, Qian-Wei Wang, Xudong Ren 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual InversionRinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik 等ICLR 2023 · 被引用 464 次
- Probabilistic Face EmbeddingsYichun Shi, Anil K. JainICCV 2019 · 被引用 362 次
相关 Paper
- Heterogeneous Feature Fusion and Cross-modal Alignment for Composed Image RetrievalGangjian Zhang, Shikui Wei, Huaxin Pang, Yao ZhaoACM MM 2021 · 被引用 34 次
- Towards Robust Uncertainty Calibration for Composed Image RetrievalYifan Wang, Wuliang Huang, Yufan Wen, Shunning Liu 等NeurIPS 2025
- HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video RetrievalZhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu 等ACM MM 2025 · 被引用 5 次
- Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty RegularizationYiyang Chen, Zhedong Zheng, Wei Ji, Leigang Qu 等ICLR 2024 · 被引用 80 次
- ConText-CIR: Learning from Concepts in Text for Composed Image RetrievalEric Xing, Pranavi Kolouju, Robert Pless, Abby Stylianou 等CVPR 2025
