EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models
Jincheng Xie, Xingchen Xiao, Runheng Liu, Zhongyi Huang, Yu Zheng, Heyan Huang
摘要
Unified multimodal embedding spaces underpin practical applications such as cross-modal retrieval and zero-shot recognition. In many real deployments, however, supervision is available only for a small subset of modality pairs (e.g., image—text), leaving unpaired modality pairs (e.g., audio↔depth, infrared↔audio) weakly connected and thus performing poorly on zero-shot transfer. Addressing this sparse-pairing regime is therefore essential for scaling unified embedding systems to new tasks without curating exhaustive pairwise data. We propose EmergentBridge, an embedding-level bridging framework that improves performance on these unpaired pairs without requiring exhaustive pairwise supervision. Our key observation is that naively aligning a new modality to a synthesized proxy embedding can introduce gradient interference, degrading the anchor-alignment structure that existing retrieval/classification relies on. EmergentBridge addresses this by (i) learning a mapping that produces a noisy bridge anchor (a proxy embedding of an already-aligned modality) from an anchor embedding, and (ii) enforcing proxy alignment only in the subspace orthogonal to the anchor-alignment direction, preserving anchor alignment while strengthening non-anchor connectivity. Across nine datasets spanning multiple modalities, EmergentBridge consistently outperforms prior binding baselines on zero-shot classification and retrieval, demonstrating strong emergent alignment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung 等NeurIPS 2022 · 被引用 834 次
- Self-Supervised MultiModal Versatile NetworksJean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic 等NeurIPS 2020 · 被引用 423 次
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic AlignmentBin Zhu, Bin Lin, Munan Ning, Yang Yan 等ICLR 2024 · 被引用 403 次
相关 Paper
- ImageBind One Embedding Space to Bind Them AllRohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh 等CVPR 2023
- Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion BridgesEitan Kosman, Gabriele Serussi, Chaim BaskinICML 2026
- TextME: Bridging Unseen Modalities Through Text DescriptionsSoyeon Hong, Jinchan Kim, Jaegook You, Seungtaek Choi 等ICML 2026
- Learning Cross-Aligned Latent Embeddings for Zero-Shot Cross-Modal RetrievalKaiyi Lin, Xing Xu, Lianli Gao, Zheng Wang 等AAAI 2020 · 被引用 50 次
- TAViS: Text-bridged Audio-Visual Segmentation with Foundation ModelsZiyang Luo, Nian Liu, Xuguang Yang, Salman Khan 等ICCV 2025
