Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment
Yang Liu, Mengyuan Liu, Shudong Huang, Jiancheng Lv
摘要
Learning visual semantic similarity is a critical challenge in bridging the gap between images and texts. However, there exist inherent variations between vision and language data, such as information density, i.e., images can contain textual information from multiple different views, which makes it difficult to compute the similarity between these two modalities accurately and efficiently. In this paper, we propose a novel framework called Asymmetric Visual Semantic Embedding (AVSE) to dynamically select features from various regions of images tailored to different textual inputs for similarity calculation. To capture information from different views in the image, we design a radial bias sampling module to sample image patches and obtain image features from various views, Furthermore, AVSE introduces a novel module for efficient computation of visual semantic similarity between asymmetric image and text embeddings. Central to this module is the presumption of foundational semantic units within the embeddings, denoted as ``meta-semantic embeddings." It segments all embeddings into meta-semantic embeddings with the same dimension and calculates visual semantic similarity by finding the optimal match of meta-semantic embeddings of two modalities. Our proposed AVSE model is extensively evaluated on the large-scale MS-COCO and Flickr30K datasets, demonstrating its superiority over recent state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text MatchingYang Liu, Wentao Feng, Zhuoyao Liu, Shudong Huang 等ICCV 2025 · 被引用 3 次
- StructXLIP: Enhancing Vision-language Models with Multimodal Structural CuesZanxi Ruan, Songqun Gao, Qiuyu Kong, Yiming Wang 等CVPR 2026 · 被引用 1 次
- SEPS: Semantic-Enhanced Patch Slimming Framework for Fine-Grained Cross-Modal AlignmentXinyu Mao, Junsi Li, Haoji Zhang, Yu Liang 等ICML 2026
- Multi-View Differential Mixing and Graph-Guided Structural Region Selection for Cross-Modal AlignmentLinlin Ji, Li LiuAAAI 2026
- Intra-Modal Neighbors Never Lie: Rectifying Inter-Modal Noisy Correspondence via Graph-Based Intra-Modal ReasoningYang Liu, Wentao Feng, Shudong Huang, Yalan Ye 等ICML 2026
它引用的顶会 Paper14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li 等ICCV 2019 · 被引用 598 次
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 被引用 413 次
- Deep Evidential Learning with Noisy Correspondence for Cross-modal RetrievalYang Qin, Dezhong Peng, Xi Peng, Xu Wang 等ACM MM 2022 · 被引用 101 次
相关 Paper
- Learning Semantic Relationship among Instances for Image-Text MatchingZheren Fu, Zhendong Mao, Yan Song, Yongdong ZhangCVPR 2023
- Noisy Correspondence Rectification via Asymmetric Similarity LearningYunbo Wang, YuJie Wu, Zhien Dai, Can Tian 等AAAI 2025 · 被引用 4 次
- CoV-Align: Efficient Fine-grained Cross-Modal Alignment with Cohesive Visual Semantics PriorityHengqi Liu, Wanting Zhou, Longteng Kong, Fangxiang Feng 等CVPR 2026
- Adaptive Prompt-Based Semantic Embedding with Inspire Potential of Implicit Knowledge for Cross-Modal RetrievalXin Huang, Shilong Wang, Tong Jia, Zhihang Gou 等AAAI 2025 · 被引用 2 次
- Saliency-Guided Attention Network for Image-Sentence MatchingZhong Ji, Haoran Wang, Jungong Han, Yanwei PangICCV 2019 · 被引用 96 次
