Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities
Yiyuan Zhang, Xiaohan Ding, Kaixiong Gong, Yixiao Ge, Ying Shan, Xiangyu Yue
Abstract
We propose to improve transformers of a specific modality with irrelevant data from other modalities, e.g., improve an ImageNet model with audio or point cloud datasets. We would like to highlight that the data samples of the target modality are irrelevant to the other modalities, which distinguishes our method from other works utilizing paired (e.g., CLIP) or interleaved data of different modalities. We propose a methodology named Multimodal Pathway - given a target modality and a transformer designed for it, we use an auxiliary transformer trained with data of another modality and construct pathways to connect components of the two models so that data of the target modality can be processed by both models. In this way, we utilize the universal sequence-to-sequence modeling abilities of transformers obtained from two modalities. As a concrete implementation, we use a modality-specific tokenizer and task-specific head as usual but utilize the transformer blocks of the auxiliary model via a proposed method named Cross-Modal Re-parameterization, which exploits the auxiliary weights without any inference costs. On the image, point cloud, video, and audio recognition tasks, we observe significant and consistent performance improvements with irrelevant data from other modalities. The code and models are available at https://github.com/AILab-CVC/M2PT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e25fa12e-7f9c-445b-a2fa-0c6fd4eee2d9Cited by top-tier papers3
- DecompNet: Enhancing Time Series Forecasting Models with Implicit DecompositionDonghao Luo, Xue WangNeurIPS 2025 · 2 citations
- UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image RecognitionXiaohan Ding, Yiyuan Zhang, Yixiao Ge, Sijie Zhao et al.CVPR 2024
- ABC-Former: Auxiliary Bimodal Cross-domain Transformer with Interactive Channel Attention for White BalanceYu-Cheng Chiu, Guan-Rong Chen, Zihao Chen, Yan-Tsung PengCVPR 2025
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal ModelsSharut Gupta, Shobhita Sundaram, Chenyu Wang, Stefanie Jegelka et al.ICLR 2026
- OmniVec2 - A Novel Transformer Based Network for Large Scale Multimodal and Multitask LearningSiddharth Srivastava, Gaurav SharmaCVPR 2024 · 32 citations
- Non-Linguistic Supervision for Contrastive Learning of Sentence EmbeddingsYiren Jian, Chongyang Gao, Soroush VosoughiNeurIPS 2022 · 20 citations
- 2D-3D Interlaced Transformer for Point Cloud Segmentation with Scene-Level SupervisionCheng-Kun Yang, Min-Hung Chen, Yung-Yu Chuang, Yen-Yu LinICCV 2023 · 30 citations
- Learnable Irrelevant Modality Dropout for Multimodal Action Recognition on Modality-Specific Annotated VideosSaghir Alfasly, Jian Lu, Chen Xu, Yuru ZouCVPR 2022 · 28 citations
