MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality Assessment
Huangbiao Xu, Huanqi Wu, Xiao Ke, Junyi Wu, Rui Xu, Jinglin Xu
Abstract
Multimodal Action Quality Assessment (AQA) has recently emerged as a promising paradigm. By leveraging complementary information across shared contextual cues, it enhances the discriminative evaluation of subtle intra-class variations in highly similar action sequences. However, partial modalities are frequently unavailable at the inference stage in reality. The absence of any modality often renders existing multimodal models inoperable. Furthermore, it triggers catastrophic performance degradation due to interruptions in cross-modal interactions. To address this issue, we propose a novel Missing Completion Framework with Mixture of Experts (MCMoE) that unifies unimodal and joint representation learning in single-stage training. Specifically, we propose an adaptive gated modality generator that dynamically fuses available information to reconstruct missing modalities. We then design modality experts to learn unimodal knowledge and dynamically mix the knowledge of all experts to extract cross-modal joint representations. With a mixture of experts, missing modalities are further refined and complemented. Finally, in the training phase, we mine the complete multimodal features and unimodal expert knowledge to guide modality generation and generation-based joint representation extraction. Extensive experiments demonstrate that our MCMoE achieves state-of-the-art results in both complete and incomplete multimodal learning on three public AQA benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 729a71ef-4118-4853-b783-71b05b96cf7cCited by top-tier papers1
Ask how each one uses itBuilds on20
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- Incomplete Multimodality-Diffused Emotion RecognitionYuanzhi Wang, Yong Li, Zhen CuiNeurIPS 2023 · 155 citations
- Action Assessment by Joint Relation GraphsJiahui Pan, Jibin Gao, Wei-Shi ZhengICCV 2019 · 141 citations
- FineDiving: A Fine-grained Dataset for Procedure-aware Action Quality AssessmentJinglin Xu, Yongming Rao, Xumin Yu, Guangyi Chen et al.CVPR 2022 · 118 citations
Related papers
- Taming Cascaded Mixture-of-Experts for Modality-missing Multi-modal Salient Object DetectionKunpeng Wang, Feifan Sun, Keke ChenAAAI 2026
- Leveraging Knowledge of Modality Experts for Incomplete Multimodal LearningWenxin Xu, Hexin Jiang, Xuefeng LiangACM MM 2024 · 31 citations
- Multimodal Emotion Recognition with Missing Modality via a Unified Multi-task Pre-training FrameworkZiyi Li, Wei-Long Zheng, Bao-Liang LuACM MM 2025 · 2 citations
- QA-MoE: Towards a Continuous Reliability Spectrum with Quality-Aware Mixture of Experts for Robust Multimodal Sentiment AnalysisYitong Zhu, Yuxuan Jiang, Guanxuan Jiang, Bojing Hou et al.ACL 2026
- FuseMoE: Mixture-of-Experts Transformers for Fleximodal FusionXing Han, Huy Nguyen, Carl Harris, Nhat Ho et al.NeurIPS 2024 · 129 citations
