Dynamic Modality Interaction Modeling for Image-Text Retrieval
Leigang Qu, Meng Liu, Jianlong Wu, Zan Gao, Liqiang Nie
摘要
Image-text retrieval is a fundamental and crucial branch in information retrieval. Although much progress has been made in bridging vision and language, it remains challenging because of the difficult intra-modal reasoning and cross-modal alignment. Existing modality interaction methods have achieved impressive results on public datasets. However, they heavily rely on expert experience and empirical feedback towards the design of interaction patterns, therefore, lacking flexibility. To address these issues, we develop a novel modality interaction modeling network based upon the routing mechanism, which is the first unified and dynamic multimodal interaction framework towards image-text retrieval. In particular, we first design four types of cells as basic units to explore different levels of modality interactions, and then connect them in a dense strategy to construct a routing space. To endow the model with the capability of path decision, we integrate a dynamic router in each cell for pattern exploration. As the routers are conditioned on inputs, our model can dynamically learn different activated paths for different data. Extensive experiments on two benchmark datasets, i.e., Flickr30K and MS-COCO, verify the superiority of our model compared with several state-of-the-art baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- Negative-Aware Attention Framework for Image-Text MatchingKun Zhang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2022 · 被引用 185 次
- LayoutLLM-T2I: Eliciting Layout Guidance from LLM for Text-to-Image GenerationLeigang Qu, Shengqiong Wu, Hao Fei, Liqiang Nie 等ACM MM 2023 · 被引用 91 次
- Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty RegularizationYiyang Chen, Zhedong Zheng, Wei Ji, Leigang Qu 等ICLR 2024 · 被引用 80 次
- Your Negative May not Be True Negative: Boosting Image-Text Matching with False Negative EliminationHaoxuan Li, Yi Bin, Junrong Liao, Yang Yang 等ACM MM 2023 · 被引用 42 次
- Dynamic Routing Transformer Network for Multimodal Sarcasm DetectionYuan Tian, Nan Xu, Ruike Zhang, Wenji MaoACL 2023 · 被引用 40 次
它引用的顶会 Paper11
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li 等ICCV 2019 · 被引用 598 次
- CAMP: Cross-Modal Adaptive Message Passing for Text-Image RetrievalZihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng 等ICCV 2019 · 被引用 349 次
- Context-Aware Multi-View Summarization Network for Image-Text MatchingLeigang Qu, Meng Liu, Da Cao, Liqiang Nie 等ACM MM 2020 · 被引用 159 次
- Adaptive Cross-Modal Embeddings for Image-Text AlignmentJonatas Wehrmann, Camila Kolling, Rodrigo C. BarrosAAAI 2020 · 被引用 86 次
- Expressing Objects Just Like Words: Recurrent Visual Embedding for Image-Text MatchingTianlang Chen, Jiebo LuoAAAI 2020 · 被引用 71 次
相关 Paper
- D2R: Dual-Branch Dynamic Routing Network for Multimodal Sentiment DetectionYifan Chen, Kuntao Li, Weixing Mai, Qiaofeng Wu 等EMNLP 2024 · 被引用 10 次
- External Knowledge Dynamic Modeling for Image-text RetrievalSong Yang, Qiang Li, Wenhui Li, Min Liu 等ACM MM 2023 · 被引用 1 次
- Learnable Pillar-based Re-ranking for Image-Text RetrievalLeigang Qu, Meng Liu, Wenjie Wang, Zhedong Zheng 等SIGIR 2023 · 被引用 23 次
- Giving Text More Imagination Space for Image-text MatchingXinfeng Dong, Longfei Han, Dingwen Zhang, Li Liu 等ACM MM 2023 · 被引用 9 次
- Improving Fusion of Region Features and Grid Features via Two-Step Interaction for Image-Text RetrievalDongqing Wu, Huihui Li, Cang Gu, Lei Guo 等ACM MM 2022 · 被引用 10 次
