Preserving Modality Structure Improves Multi-Modal Learning
Sirnam Swetha, Mamshad Nayeem Rizve, Nina Shvetsova, Hilde Kuehne, Mubarak Shah
Abstract
Self-supervised learning on large-scale multi-modal datasets allows learning semantically meaningful embeddings in a joint multi-modal representation space without relying on human annotations. These joint embeddings enable zero-shot cross-modal tasks like retrieval and classification. However, these methods often struggle to generalize well on out-of-domain data as they ignore the semantic structure present in modality-specific embeddings. In this context, we propose a novel Semantic-Structure-Preserving Consistency approach to improve generalizability by preserving the modality-specific relationships in the joint embedding space. To capture modality-specific semantic relationships between samples, we propose to learn multiple anchors and represent the multifaceted relationship between samples with respect to their relationship with these anchors. To assign multiple anchors to each sample, we propose a novel Multi-Assignment Sinkhorn-Knopp algorithm. Our experimentation demonstrates that our proposed approach learns semantically meaningful anchors in a self-supervised manner. Furthermore, our evaluation on MSR-VTT and YouCook2 datasets demonstrates that our proposed multi-anchor assignment based solution achieves state-of-the-art performance and generalizes to both in-and out-of-domain datasets. Code: https://github.com/Swetha5/Multi_Sinkhorn_Knopp
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Learning Unseen Modality InteractionYunhua Zhang, Hazel Doughty, Cees SnoekNeurIPS 2023 · 16 citations
- Structured Spectral Reasoning for Frequency-Adaptive Multimodal RecommendationWei Yang, Rui Zhong, Yiqun Chen, Chi Lu et al.NeurIPS 2025 · 10 citations
- TIGER: A Unified Framework for Time, Images and Geo-location RetrievalDavid G. Shatwell, Sirnam Swetha, Mubarak ShahCVPR 2026 · 2 citations
- Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video UnderstandingJoseph Fioresi, Ishan Rajendrakumar Dave, Mubarak ShahICLR 2026 · 1 citation
- GT-Loc: Unifying When and Where in Images Through a Joint Embedding SpaceDavid G. Shatwell, Ishan Rajendrakumar Dave, Sirnam Swetha, Mubarak ShahICCV 2025 · 1 citation
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 873 citations
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani et al.NeurIPS 2020 · 483 citations
Related papers
- Multimodal Clustering Networks for Self-supervised Learning from Unlabeled VideosBrian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne et al.ICCV 2021 · 98 citations
- Normalized Contrastive Learning for Text-Video RetrievalYookoon Park, Mahmoud Azab, Seungwhan Moon, Bo Xiong et al.EMNLP 2022 · 10 citations
- Multi-Modal Self-Supervised Learning for RecommendationWei Wei, Chao Huang, Lianghao Xia, Chuxu ZhangWWW 2023 · 256 citations
- AG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal ModelsZhen Qu, Xian Tao, Xiaoyi Bao, Dingrong Wang et al.CVPR 2026 · 1 citation
- Mitigating Noisy Correspondence by Geometrical Structure Consistency LearningZihua Zhao, Mengxi Chen, Tianjie Dai, Jiangchao Yao et al.CVPR 2024 · 5 citations
