Toward Explainable Physical Audiovisual Commonsense Reasoning
Daoming Zong, Chaoyue Ding, Kaitao Chen
Abstract
For AI systems to be safely and reliably grounded in the real world, they should possess the ability of physical commonsense reasoning. Physical commonsense reasoning is essentially a multisensory task as physical properties of objects are manifested through multiple perception modalities, including both visual and auditory. In this study, we constructed two new benchmarks, called PACS-Reason and PACS-Reason+, for explainable physical audiovisual commonsense reasoning (EPACS), in which each datapoint is accompanied by a golden detailed rationale (intermediate reasoning path) to explain the answer selection. Moreover, we present PAVC-Reasoner, a multimodal large language model (LLM) designed to reason about physical commonsense attributes. The model aligns different modalities with the language modality by integrating three different perceivers for cross-modal pretraining and instruction finetuning at multiple granularities. It utilizes an LLM as a cognitive engine to process multimodal inputs and output convincing intermediate reasoning paths as justification for inferring answers. Numerous experiments have demonstrated the effectiveness and superiority of PAVC-Reasoner as a baseline model for studying EPACS. Most attractively, PAVC-Reasoner is capable of reasoning and obtaining strong interpretable explicit reasoning paths, signifying a significant stride towards real-world physical commonsense reasoning.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 07fa20be-ff09-436c-8f43-dbedbaf9624cCited by top-tier papers1
Ask how each one uses itRelated papers
- McOmet: Multimodal Fusion Transformer for Physical Audiovisual Commonsense ReasoningDaoming Zong, Shiliang SunAAAI 2023 · 5 citations
- PAI-Bench: A Comprehensive Benchmark For Physical AIFengzhe Zhou, Jiannan Huang, Jialuo Li, Deva Ramanan et al.CVPR 2026 · 32 citations
- Counterfactual Debiasing for Physical Audiovisual Commonsense ReasoningDaoming Zong, Chaoyue Ding, Kaitao Chen, Yinsheng Li et al.AAAI 2025
- SpaCE-Eval: A Benchmark for Real-World Multi-Modal ReasoningXuyou Yang, Yucheng Zhao, Wenxuan Zhang, Immanuel KohICLR 2026
- VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual ReasoningHao Yan, Xingchen Liu, Hao Wang, Zhenbiao Cao et al.ICLR 2026 · 7 citations
