MuMu: Cooperative Multitask Learning-Based Guided Multimodal Fusion
Md Mofijul Islam, Tariq Iqbal
Abstract
Multimodal sensors (visual, non-visual, and wearable) can provide complementary information to develop robust perception systems for recognizing activities accurately. However, it is challenging to extract robust multimodal representations due to the heterogeneous characteristics of data from multimodal sensors and disparate human activities, especially in the presence of noisy and misaligned sensor data. In this work, we propose a cooperative multitask learningbased guided multimodal fusion approach, MuMu, to extract robust multimodal representations for human activity recognition (HAR). MuMu employs an auxiliary task learning approach to extract features specific to each set of activities with shared characteristics (activity-group). MuMu then utilizes activity-group-specific features to direct our proposed Guided Multimodal Fusion Approach (GM-Fusion) for extracting complementary multimodal representations, designed as the target task. We evaluated MuMu by comparing its performance to state-of-the-art multimodal HAR approaches on three activity datasets. Our extensive experimental results suggest that MuMu outperforms all the evaluated approaches across all three datasets. Additionally, the ablation study suggests that MuMu significantly outperforms the baseline models (p < 0.05), which do not use our guided multimodal fusion. Finally, the robust performance of MuMu on noisy and misaligned sensor data posits that our approach is suitable for HAR in real-world settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca278a69-7f34-4622-809d-22afd73da4beCited by top-tier papers7
- WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity RecognitionMarius Bock, Hilde Kuehne, Kristof Van Laerhoven, Michael MöllerUbiComp 2025 · 50 citations
- MMTSA: Multi-Modal Temporal Segment Attention Network for Efficient Human Activity RecognitionZiqi Gao, Yuntao Wang, Jianguo Chen, Junliang Xing et al.UbiComp 2023 · 22 citations
- EQA-MX: Embodied Question Answering using Multimodal ExpressionMd Mofijul Islam, Alexi Gladstone, Riashat Islam, Tariq IqbalICLR 2024 · 18 citations
- PATRON: Perspective-Aware Multitask Model for Referring Expression Grounding Using Embodied Multimodal CuesMd Mofijul Islam, Alexi Gladstone, Tariq IqbalAAAI 2023 · 10 citations
- FAMOS: Robust Privacy-Preserving Authentication on Payment Apps via Federated Multi-Modal Contrastive LearningYifeng Cai, Ziqi Zhang, Jiaping Gui, Bingyan Liu et al.USENIX Security 2024 · 6 citations
Builds on8
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Task2Vec: Task Embedding for Meta-LearningAlessandro Achille, Michael Lam, Rahul Tewari, Avinash Ravichandran et al.ICCV 2019 · 359 citations
- Learning to Branch for Multi-Task LearningPengsheng Guo, Chen-Yu Lee, Daniel UlbrichtICML 2020 · 208 citations
- MMAct: A Large-Scale Dataset for Cross Modal Human Action UnderstandingQuan Kong, Ziming Wu, Ziwei Deng, Martin Klinkigt et al.ICCV 2019 · 108 citations
- DeepTake: Prediction of Driver Takeover Behavior using Multimodal DataErfan Pakdamanian, Shili Sheng, Sonia Baee, Seongkook Heo et al.CHI 2021 · 86 citations
Related papers
- MASTER: A Multi-modal Foundation Model for Human Activity RecognitionGuanzhou Zhu, Dong Zhao, Chunliang Li, Mingyue Zhao et al.UbiComp 2025 · 8 citations
- Adversarial Multi-view Networks for Activity RecognitionLei Bai, Lina Yao, Xianzhi Wang, Salil S. Kanhere et al.UbiComp 2020 · 41 citations
- Augmented Adversarial Learning for Human Activity Recognition with Partial Sensor SetsHua Kang, Qianyi Huang, Qian ZhangUbiComp 2022 · 16 citations
- METIER: A Deep Multi-Task Learning Based Activity and User Recognition Model Using Wearable SensorsLing Chen, Yi Zhang, Liangying PengUbiComp 2020 · 71 citations
- ColloSSL: Collaborative Self-Supervised Learning for Human Activity RecognitionYash Jain, Chi Ian Tang, Chulhong Min, Fahim Kawsar et al.UbiComp 2022 · 113 citations
