Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception
Hassan Akbari, Dan Kondratyuk, Yin Cui, Rachel Hornung, Huisheng Wang, Hartwig Adam
Abstract
We present Integrated Multimodal Perception (IMP), a simple and scalable multimodal multi-task training and modeling approach. IMP integrates multimodal inputs including image, video, text, and audio into a single Transformer encoder with minimal modality-specific components. IMP makes use of a novel design that combines Alternating Gradient Descent (AGD) and Mixture-of-Experts (MoE) for efficient model and task scaling. We conduct extensive empirical studies and reveal the following key insights: 1) Performing gradient descent updates by alternating on diverse modalities, loss functions, and tasks, with varying input resolutions, efficiently improves the model. 2) Sparsification with MoE on a single modality-agnostic encoder substantially improves the performance, outperforming dense models that use modality-specific encoders or additional fusion layers and greatly mitigates the conflicts between modalities. IMP achieves competitive performance on a wide range of downstream tasks including video classification, image classification, image-text, and video-text retrieval. Most notably, we train a sparse IMP-MoE-L variant focusing on video tasks that achieves new state-of-the-art in zero-shot video classification: 77.0% on Kinetics-400, 76.8% on Kinetics-600, and 68.3% on Kinetics-700, improving the previous state-of-the-art by +5%, +6.7%, and +5.8%, respectively, while using only 15% of their total training computational cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 756edf6c-4f82-498c-b149-511d1ae3bc2dCited by top-tier papers10
- VideoPoet: A Large Language Model for Zero-Shot Video GenerationDan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama et al.ICML 2024 · 464 citations
- VideoPrism: A Foundational Visual Encoder for Video UnderstandingLong Zhao, Nitesh Bharadwaj Gundavarapu, Liangzhe Yuan, Hao Zhou et al.ICML 2024 · 91 citations
- Mixtures of Experts for Audio-Visual LearningYing Cheng, Yang Li, Junjie He, Rui FengNeurIPS 2024 · 22 citations
- XTrack: Multimodal Training Boosts RGB-X Video Object TrackersYuedong Tan, Zongwei Wu, Yuqian Fu, Zhuyun Zhou et al.ICCV 2025 · 10 citations
- Mamba-CAD: State Space Model for 3D Computer-Aided Design Generative ModelingXueyang Li, Yunzhong Lou, Yu Song, Xiangdong ZhouAAAI 2025 · 6 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of ExpertsBasil Mustafa, Carlos Riquelme, Joan Puigcerver, Rodolphe Jenatton et al.NeurIPS 2022 · 359 citations
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang et al.NeurIPS 2021 · 782 citations
- UNIALIGN: Scaling Multimodal Alignment within One Unified ModelBo Zhou, Liulei Li, Yujia Wang, Huafeng Liu et al.CVPR 2025
- Everything at Once - Multi-modal Fusion Transformer for Video RetrievalNina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas et al.CVPR 2022 · 4 citations
- OmniVec2 - A Novel Transformer Based Network for Large Scale Multimodal and Multitask LearningSiddharth Srivastava, Gaurav SharmaCVPR 2024 · 32 citations
