4M: Massively Multimodal Masked Modeling
David Mizrahi, Roman Bachmann, Oguzhan Fatih Kar, Teresa Yeo, Mingfei Gao, Afshin Dehghan, Amir Zamir
摘要
Current machine learning models for vision are often highly specialized and limited to a single modality and task. In contrast, recent large language models exhibit a wide range of capabilities, hinting at a possibility for similarly versatile models in computer vision. In this paper, we take a step in this direction and propose a multimodal training scheme called 4M. It consists of training a single unified Transformer encoder-decoder using a masked modeling objective across a wide range of input/output modalities -including text, images, geometric, and semantic modalities, as well as neural network feature maps. 4M achieves scalability by unifying the representation space of all modalities through mapping them into discrete tokens and performing multimodal masked modeling on a small randomized subset of tokens. 4M leads to models that exhibit several key capabilities: (1) they can perform a diverse set of vision tasks out of the box, (2) they excel when fine-tuned for unseen downstream tasks or new input modalities, and (3) they can function as a generative model that can be conditioned on arbitrary modalities, enabling a wide variety of expressive multimodal editing capabilities with remarkable flexibility. Through experimental analyses, we demonstrate the potential of 4M for training versatile and scalable foundation models for vision tasks, setting the stage for further exploration in multimodal learning for vision and other domains. † For clarity, "modalities" usually denote the inputs to a model (e.g. sensory signals), and "tasks" usually denote the outputs (e.g. semantics). Our method enables a symmetric input-output structure, thus we use "modalities" and "tasks" interchangeably in this paper. * Equal contribution & corresponding authors. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Boosting Generative Image Modeling via Joint Image-Feature SynthesisTheodoros Kouzelis, Efstathios Karypidis, Ioannis Kakogeorgiou, Spyridon Gidaris 等NeurIPS 2025 · 被引用 47 次
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual TasksCanyu Zhao, Yanlong Sun, Mingyu Liu, Huanyi Zheng 等NeurIPS 2025 · 被引用 45 次
- Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal InputsMustafa Shukor, Matthieu CordNeurIPS 2024 · 被引用 27 次
- Adapting Diffusion Models for Improved Prompt Compliance and Controllable Image SynthesisDeepak Sridhar, Abhishek Peri, Rohith Rachala, Nuno VasconcelosNeurIPS 2024 · 被引用 5 次
- OmniResponse: Online Multimodal Conversational Response Generation in Dyadic InteractionsCheng Luo, Jianghui Wang, Bing Li, Siyang Song 等NeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper62
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
相关 Paper
- 4M-21: An Any-to-Any Vision Model for Tens of Tasks and ModalitiesRoman Bachmann, Oguzhan Fatih Kar, David Mizrahi, Ali Garjani 等NeurIPS 2024 · 被引用 60 次
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 被引用 354 次
- EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoEJunyi Chen, Longteng Guo, Jia Sun, Shuai Shao 等AAAI 2024 · 被引用 25 次
- Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language TasksWenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck 等CVPR 2023
- MIO: A Foundation Model on Multimodal TokensZekun Moore Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou 等EMNLP 2025
