Preserving Modality Structure Improves Multi-Modal Learning
Sirnam Swetha, Mamshad Nayeem Rizve, Nina Shvetsova, Hilde Kuehne, Mubarak Shah
摘要
Self-supervised learning on large-scale multi-modal datasets allows learning semantically meaningful embeddings in a joint multi-modal representation space without relying on human annotations. These joint embeddings enable zero-shot cross-modal tasks like retrieval and classification. However, these methods often struggle to generalize well on out-of-domain data as they ignore the semantic structure present in modality-specific embeddings. In this context, we propose a novel Semantic-Structure-Preserving Consistency approach to improve generalizability by preserving the modality-specific relationships in the joint embedding space. To capture modality-specific semantic relationships between samples, we propose to learn multiple anchors and represent the multifaceted relationship between samples with respect to their relationship with these anchors. To assign multiple anchors to each sample, we propose a novel Multi-Assignment Sinkhorn-Knopp algorithm. Our experimentation demonstrates that our proposed approach learns semantically meaningful anchors in a self-supervised manner. Furthermore, our evaluation on MSR-VTT and YouCook2 datasets demonstrates that our proposed multi-anchor assignment based solution achieves state-of-the-art performance and generalizes to both in-and out-of-domain datasets. Code: https://github.com/Swetha5/Multi_Sinkhorn_Knopp
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Learning Unseen Modality InteractionYunhua Zhang, Hazel Doughty, Cees SnoekNeurIPS 2023 · 被引用 16 次
- Structured Spectral Reasoning for Frequency-Adaptive Multimodal RecommendationWei Yang, Rui Zhong, Yiqun Chen, Chi Lu 等NeurIPS 2025 · 被引用 10 次
- TIGER: A Unified Framework for Time, Images and Geo-location RetrievalDavid G. Shatwell, Sirnam Swetha, Mubarak ShahCVPR 2026 · 被引用 2 次
- Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video UnderstandingJoseph Fioresi, Ishan Rajendrakumar Dave, Mubarak ShahICLR 2026 · 被引用 1 次
- GT-Loc: Unifying When and Where in Images Through a Joint Embedding SpaceDavid G. Shatwell, Ishan Rajendrakumar Dave, Sirnam Swetha, Mubarak ShahICCV 2025 · 被引用 1 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- Self-labelling via simultaneous clustering and representation learningYuki Markus Asano, Christian Rupprecht, Andrea VedaldiICLR 2020 · 被引用 873 次
- Self-Supervised Learning by Cross-Modal Audio-Video ClusteringHumam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani 等NeurIPS 2020 · 被引用 483 次
相关 Paper
- Multimodal Clustering Networks for Self-supervised Learning from Unlabeled VideosBrian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne 等ICCV 2021 · 被引用 98 次
- Normalized Contrastive Learning for Text-Video RetrievalYookoon Park, Mahmoud Azab, Seungwhan Moon, Bo Xiong 等EMNLP 2022 · 被引用 10 次
- Multi-Modal Self-Supervised Learning for RecommendationWei Wei, Chao Huang, Lianghao Xia, Chuxu ZhangWWW 2023 · 被引用 256 次
- AG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal ModelsZhen Qu, Xian Tao, Xiaoyi Bao, Dingrong Wang 等CVPR 2026 · 被引用 1 次
- Mitigating Noisy Correspondence by Geometrical Structure Consistency LearningZihua Zhao, Mengxi Chen, Tianjie Dai, Jiangchao Yao 等CVPR 2024 · 被引用 5 次
