Learning the Best Pooling Strategy for Visual Semantic Embedding
Jiacheng Chen, Hexiang Hu, Hao Wu, Yuning Jiang, Changhu Wang
摘要
Visual Semantic Embedding (VSE) is a dominant approach for vision-language retrieval, which aims at learning a deep embedding space such that visual data are embedded close to their semantic text labels or descriptions. Recent VSE models use complex methods to better contextualize and aggregate multi-modal features into holistic embeddings. However, we discover that surprisingly simple (but carefully selected) global pooling functions (e.g., max pooling) outperform those complex models, across different feature extractors. Despite its simplicity and effectiveness, seeking the best pooling function for different data modality and feature extractor is costly and tedious, especially when the size of features varies (e.g., text, video). Therefore, we propose a Generalized Pooling Operator (GPO), which learns to automatically adapt itself to the best pooling strategy for different features, requiring no manual tuning while staying effective and efficient. We extend the VSE model using this proposed GPO and denote it as VSE∞. Without bells and whistles, VSE∞ outperforms previous VSE methods significantly on image-text retrieval benchmarks across popular feature extractors. With a simple adaptation, variants of VSE∞ further demonstrate its strength by achieving the new state of the art on two video-text retrieval datasets. Comprehensive experiments and visualizations confirm that GPO always discovers the best pooling strategy and can be a plug-and-play feature aggregation module for standard VSE models. Code and pre-trained models are available at http://jcchen.me/vse_infty/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper94
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 被引用 907 次
- LiT: Zero-Shot Transfer with Locked-image text TuningXiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner 等CVPR 2022 · 被引用 349 次
- Negative-Aware Attention Framework for Image-Text MatchingKun Zhang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2022 · 被引用 185 次
- Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia EntitiesHexiang Hu, Yi Luan, Yang Chen, Urvashi Khandelwal 等ICCV 2023 · 被引用 123 次
它引用的顶会 Paper9
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li 等ICCV 2019 · 被引用 598 次
- FSPool: Learning Set Representations with Featurewise Sort PoolingYan Zhang, Jonathon S. Hare, Adam Prügel-BennettICLR 2020 · 被引用 92 次
- ACMM: Aligned Cross-Modal Memory for Few-Shot Image and Sentence MatchingYan Huang, Liang WangICCV 2019 · 被引用 68 次
相关 Paper
- VISTA: Visualized Text Embedding For Universal Multi-Modal RetrievalJunjie Zhou, Zheng Liu, Shitao Xiao, Bo Zhao 等ACL 2024
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-languageDelong Chen, Mustafa Shukor, Théo Moutakanni, Willy Chung 等ICLR 2026 · 被引用 60 次
- X-Pool: Cross-Modal Language-Video Attention for Text-Video RetrievalSatya Krishna Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan 等CVPR 2022 · 被引用 190 次
- Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language AlignmentYang Liu, Mengyuan Liu, Shudong Huang, Jiancheng LvAAAI 2025 · 被引用 8 次
- Unifying Multimodal Retrieval via Document Screenshot EmbeddingXueguang Ma, Sheng-Chieh Lin, Minghan Li, Wenhu Chen 等EMNLP 2024 · 被引用 13 次
