McOmet: Multimodal Fusion Transformer for Physical Audiovisual Commonsense Reasoning
Daoming Zong, Shiliang Sun
Abstract
Physical commonsense reasoning is essential for building reliable and interpretable AI systems, which involves a general understanding of the physical properties and affordances of everyday objects, how these objects can be manipulated, and how they interact with others. It is fundamentally a multisensory task, as physical properties are manifested through multiple modalities, including vision and acoustics. In this work, we present a unified framework, named Multimodal Commonsense Transformer (MCOMET ), for physical audiovisual commonsense reasoning. MCOMET has two intriguing properties: i) it fully mines higher-ordered temporal relationships across modalities (e.g., pairs, triplets, and quadruplets); and ii) it restricts the cross-modal flow through the feature collection and propagation mechanism with tight fusion bottlenecks, forcing the model to attend the most relevant parts in each modality and suppressing the dissemination of noisy information. We evaluate our model on a very recent public benchmark, PACS. Results show that MCOMET significantly outperforms a variety of strong baselines, revealing powerful multi-modal commonsense reasoning capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6f2ee40-e8ac-477a-802c-51d19387b8cdCited by top-tier papers2
- Attention Bootstrapping for Multi-Modal Test-Time AdaptationYusheng Zhao, Junyu Luo, Xiao Luo, Jinsheng Huang et al.AAAI 2025 · 5 citations
- DIVE: Towards Descriptive and Diverse Visual Commonsense GenerationJun-Hyung Park, Hyuntae Park, Youjin Kang, Eojin Jeon et al.EMNLP 2023
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals et al.ICML 2021 · 1,399 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
Related papers
- Toward Explainable Physical Audiovisual Commonsense ReasoningDaoming Zong, Chaoyue Ding, Kaitao ChenACM MM 2024
- Counterfactual Debiasing for Physical Audiovisual Commonsense ReasoningDaoming Zong, Chaoyue Ding, Kaitao Chen, Yinsheng Li et al.AAAI 2025
- Disentangled Counterfactual Learning for Physical Audiovisual Commonsense ReasoningChangsheng Lv, Shuai Zhang, Yapeng Tian, Mengshi Qi et al.NeurIPS 2023 · 26 citations
- PAI-Bench: A Comprehensive Benchmark For Physical AIFengzhe Zhou, Jiannan Huang, Jialuo Li, Deva Ramanan et al.CVPR 2026 · 32 citations
- Detecting Violations of Physical Common Sense in Images: A Challenge Dataset and Effective ModelWeibin Wu, Zitong Wang, Zhengjie Luo, Wenqing Chen et al.ACM MM 2025
