MMG-Ego4D: Multi-Modal Generalization in Egocentric Action Recognition
Xinyu Gong, Sreyas Mohan, Naina Dhingra, Jean-Charles Bazin, Yilei Li, Zhangyang Wang, Rakesh Ranjan
Abstract
In this paper, we study a novel problem in egocentric action recognition, which we term as "Multimodal Generalization" (MMG). MMG aims to study how systems can generalize when data from certain modalities is limited or even completely missing. We thoroughly investigate MMG in the context of standard supervised action recognition and the more challenging few-shot setting for learning new action categories. MMG consists of two novel scenarios, designed to support security, and efficiency considerations in real-world applications: (1) missing modality generalization where some modalities that were present during the train time are missing during the inference time, and ( 2 ) cross-modal zero-shot generalization, where the modalities present during the inference time and the training time are disjoint. To enable this investigation, we construct a new dataset MMG-Ego4D containing data points with video, audio, and inertial motion sensor (IMU) modalities. Our dataset is derived from Ego4D [27] dataset, but processed and thoroughly re-annotated by human experts to facilitate research in the MMG problem. We evaluate a diverse array of models on MMG-Ego4D and propose new methods with improved generalization ability. In particular, we introduce a new fusion module with modality dropout training, contrastive-based alignment training, and a novel cross-modal prototypical loss for better few-shot performance. We hope this study will serve as a benchmark and guide future research in multimodal generalization problems. The benchmark and code are available at https://github.com/facebookresearch/MMG Ego4D Modality video audio IMU Memory per second of data (KB) 593.92 62.76 9.44 Typical model FLOPs (G) 70.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f357202-2836-4873-8afc-cf72b1cf86d2Cited by top-tier papers10
- EMHI: A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUsZhen Fan, Peng Dai, Zhuo Su, Xu Gao et al.AAAI 2025 · 13 citations
- MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion LearningThanh-Dat Truong, Christophe Bobda, Nitin Agarwal, Khoa LuuNeurIPS 2025 · 6 citations
- COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity RecognitionBaiyu Chen, Wilson Wongso, Zechen Li, Yonchanok Khaokaew et al.UbiComp 2026 · 1 citation
- SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View FusionXiang Li, Heqian Qiu, Lanxiao Wang, Benliu Qiu et al.CVPR 2026 · 1 citation
- RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for RoboticsChan Hee Song, Valts Blukis, Jonathan Tremblay, Stephen Tyree et al.CVPR 2025
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
Related papers
- What can a cook in Italy teach a mechanic in India? Action Recognition Generalisation Over Scenarios and LocationsChiara Plizzari, Toby Perrett, Barbara Caputo, Dima DamenICCV 2023 · 27 citations
- Test-Time Adaptation for Combating Missing Modalities in Egocentric VideosMerey Ramazanova, Alejandro Pardo, Bernard Ghanem, Motasem AlfarraICLR 2025
- Achieving Cross Modal Generalization with Multimodal Unified RepresentationYan Xia, Hai Huang, Jieming Zhu, Zhou ZhaoNeurIPS 2023 · 84 citations
- MODA: Motion-Drift Augmentation for Inertial Human Motion AnalysisYinghao Wu, Shihui Guo, Yipeng QinCVPR 2025
- Towards Good Practices for Missing Modality Robust Action RecognitionSangmin Woo, Sumin Lee, Yeonju Park, Muhammad Adi Nugroho et al.AAAI 2023 · 80 citations
