DH-Set: Improving Vision-Language Alignment with Diverse and Hybrid Set-Embeddings Learning
Kun Zhang, Jingyu Li, Zhe Li, S. Kevin Zhou
Abstract
Vision-Language (VL) alignment across image and text modalities is a challenging task due to the inherent semantic ambiguity of data with multiple possible meanings. Existing methods typically solve it by learning multiple subrepresentation spaces to encode each input data as a set of embeddings, and constraining diversity between whole subspaces to capture diverse semantics for accurate VL alignment. Despite their promising outcomes, existing methods suffer two imperfections: 1) actually, specific semantics is mainly expressed by some local dimensions within the subspace. Ignoring this intrinsic property, existing diversity constraints imposed on the whole subspace may impair diverse embedding learning; 2) multiple embeddings are inevitably introduced, sacrificing computational and storage efficiency. In this paper, we propose a simple yet effective Diverse and Hybrid Set-embeddings learning framework (DH-Set), which is distinct from prior work in three aspects. DH-Set 1) devises a novel semantic importance dissecting method to focus on key local dimensions within each subspace; and thereby 2) not only imposes finer-grained diversity constraint to improve the accuracy of diverse embedding learning, 3) but also mixes key dimensions of all subspaces into the single hybrid embedding to boost inference efficiency. Extensive experiments on various benchmarks and model backbones show the superiority of DH-Set over state-of-the-art methods, achieving substantial 2.3%-14.7% rSum improvements while lowering computational and storage complexity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80c1d2a9-448d-4579-87c6-33021cd48fc1Cited by top-tier papers3
- Hierarchy-Aware Pseudo Word Learning with Text Adaptation for Zero-Shot Composed Image RetrievalZhe Li, Lei Zhang, Zheren Fu, Kun Zhang et al.ICCV 2025 · 1 citation
- Controllable Contamination Detection for Reliable LLM Evaluation with Statistical GuaranteesZheng Zhang, Qi Liu, Siyuan Liang, Ning Li et al.ACL 2026
- Gravitation-Driven Semantic Alignment for Text Video RetrievalYi Yang, Zheng Wang, Xing Xu, Jingkuan Song et al.CVPR 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document UnderstandingAhmed Masry, Juan A. Rodríguez, Tianyu Zhang, Suyuchen Wang et al.NeurIPS 2025 · 7 citations
- HANet: Hierarchical Alignment Networks for Video-Text RetrievalPeng Wu, Xiangteng He, Mingqian Tang, Yiliang Lv et al.ACM MM 2021 · 62 citations
- Distilled Dual-Encoder Model for Vision-Language UnderstandingZekun Wang, Wenhui Wang, Haichao Zhu, Ming Liu et al.EMNLP 2022 · 22 citations
- Modality Eigen-Encodings Are Keys to Open Modality Informative ContainersYiyuan Zhang, Yuqi JiACM MM 2022
- DeepAlign: Mitigating Modality Conflict through Modality-Specific AlignmentShuo Li, Bingchen Miao, Wendong Bu, Juncheng Li et al.CVPR 2026
