BDIQA: A New Dataset for Video Question Answering to Explore Cognitive Reasoning through Theory of Mind
Yuanyuan Mao, Xin Lin, Qin Ni, Liang He
摘要
As a foundational component of cognitive intelligence, theory of mind (ToM) can make AI more closely resemble human thought processes, thereby enhancing their interaction and collaboration with human. In particular, it can significantly improve a model's comprehension of videos in complex scenes. However, current video question answer (VideoQA) datasets focus on studying causal reasoning within events, few of them genuinely incorporating human ToM. Consequently, there is a lack of development in ToM reasoning tasks within the area of VideoQA. This paper presents BDIQA, the first benchmark to explore the cognitive reasoning capabilities of VideoQA models in the context of ToM. BDIQA is inspired by the cognitive development of children's ToM and addresses the current deficiencies in machine ToM within datasets and tasks. Specifically, it offers tasks at two difficulty levels, assessing Belief, Desire and Intention (BDI) reasoning in both simple and complex scenarios. We conduct evaluations on several mainstream methods of VideoQA and diagnose their capabilities with zero-shot, few-shot and supervised learning. We find that the performance of pre-trained models on cognitive reasoning tasks remains unsatisfactory. To counter this challenge, we undertake thorough analysis and experimentation, ultimately presenting two guidelines to enhance cognitive reasoning derived from ablation analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ALLVB: All-in-One Long Video Understanding BenchmarkXichen Tan, Yuanjing Luo, Yunfan Ye, Fang Liu 等AAAI 2025 · 被引用 13 次
- ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of MindKazutoshi Shinoda, Nobukatsu Hojo, Kyosuke Nishida, Saki Mizuno 等AAAI 2025 · 被引用 10 次
- MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied AgentsRuoxuan Zhang, Qiyun Zheng, Zhiyu Zhou, Ziqi Liao 等CVPR 2026 · 被引用 6 次
它引用的顶会 Paper17
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等ICCV 2021 · 被引用 345 次
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等NeurIPS 2022 · 被引用 305 次
相关 Paper
- MMToM-QA: Multimodal Theory of Mind Question AnsweringChuanyang Jin, Yutong Wu, Jing Cao, Jiannan Xiang 等ACL 2024 · 被引用 8 次
- Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion ReasoningMeng Luo, Bobo Li, Shanqing Xu, Shize Zhang 等ICLR 2026 · 被引用 10 次
- A Very Big Video Reasoning SuiteMaijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji 等ICML 2026 · 被引用 20 次
- Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-AnsweringZhaohe Liao, Jiangtong Li, Siyu Sun, Qingyang Liu 等ICML 2025
- From Representation to Reasoning: Towards both Evidence and Commonsense Reasoning for Video Question-AnsweringJiangtong Li, Li Niu, Liqing ZhangCVPR 2022 · 被引用 48 次
