MULTIMODALITY AS SUPERVISION: SELF-SUPERVISED SPECIALIZATION TO THE TEST ENVIRONMENT VIA MULTIMODALITY
Kunal Pratap Singh, Ali Garjani, Rishubh Singh, Muhammad Uzair Khattak, Jason Toskov, Efe Tarhan, Andrei Atanov, Oguzhan Fatih Kar, Amir Zamir
Abstract
Cross-modal learning, i.e., learning to predict one modality from another, is a fundamental mechanism for self-supervision via leveraging multimodality. Many practical applications, e.g., deploying a household robot, involve devices that are equipped with a rich set of sensors that enable multimodal sensing in their test environment. This presents an opportunity to apply cross-modal learning to the multimodal data sensed by these devices to learn representations. Findings in developmental psychology also suggest that biological agents leverage it to build an effective representation of their surroundings.
To study this, we propose a sandbox, where we restrict a user device to just a given test environment. It results in a specialization setup where we attempt to develop a performant model for this specific test environment. Under this setup, we develop Test-Space Training (TST), which performs multimodal data collection in the test environment and performs self-supervised pre-training on it. We evaluate these models on various downstream tasks in the same environment. We find various interesting insights, such as collecting rich multimodal data only from the test environment and leveraging cross-modal learning, we can achieve competitive results with generalist models (Oquab et al., 2023; Radford et al., 2021), pre-trained on large-scale internet-based datasets. This enables an alternative scenario where the need for external Internet-scale datasets for pre-training models is reduced. We also present a set of analyses and ablations that raise intriguing points on substituting data with (multi)modality, and how varying pre-training data enables a tradeoff between a model’s abilities to specialise to a test environment, and generalize to held-out spaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on48
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
Related papers
- Cross-Tactile Sensor Representation LearningYan Zhang, Zheng WANG, Pengpeng Zeng, Xing Xu et al.ICML 2026
- Multimodal Clustering Networks for Self-supervised Learning from Unlabeled VideosBrian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne et al.ICCV 2021 · 98 citations
- Learning Unseen Modality InteractionYunhua Zhang, Hazel Doughty, Cees SnoekNeurIPS 2023 · 16 citations
- Shared Cross-Modal Trajectory Prediction for Autonomous DrivingChiho Choi, Joon Hee Choi, Jiachen Li, Srikanth MallaCVPR 2021
- Cross-Modal Generalization: Learning in Low Resource Modalities via Meta-AlignmentPaul Pu Liang, Peter Wu, Liu Ziyin, Louis-Philippe Morency et al.ACM MM 2021 · 22 citations
