VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
Sirnam Swetha, Rohit Gupta, Parth Parag Kulkarni, David G. Shatwell, Jeffrey A. Chan-Santiago, Nyle Siddiqui, Joseph Fioresi, Mubarak Shah
Abstract
Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly focus on questions answerable through explicit visual content - actions, objects, and events - directly observable within individual frames or short clips. To truly understand videos as humans do, models must go beyond what is directly shown, inferring hidden relationships and contextual cues that are only implied across frames. Current benchmarks fail to capture this essential aspect of video understanding. To address this gap, we introduce VRR-QA, a benchmark for Visual Relational Reasoning Beyond Explicit Cues. We curate our benchmark from creative and cinematic videos such as movies, that deliberately employ storytelling techniques which omit direct depictions of certain events or relations, requiring viewers to infer them. VRR-QA comprises 1K meticulously expert-annotated QA pairs drawn from 1K creative video clips covering 15 genres across 7 decades of content, from both live-action and animated titles. Our extensive evaluations on 14 leading VideoQA models reveals consistent and significant performance degradation, underscoring their reliance on surface-level visual cues and highlighting the difficulty of implicit reasoning. Even the best model substantially underperforms human baselines with only 64% accuracy. Performance variations across models further illustrate the complexity and diversity of the challenges presented by VRR-QA. By releasing both dataset and data collection framework, VRR-QA establishes a rigorous, diverse, and reproducible testbed for advancing VideoQA: https://swetha5.github.io/ImplicitQA/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d590a36c-58c4-41ee-bfdf-82ec834c286fBuilds on11
- IntentQA: Context-aware Video Intent ReasoningJiapeng Li, Ping Wei, Wenjuan Han, Lifeng FanICCV 2023 · 97 citations
- Scaling RL to Long VideosYukang Chen, Wei Huang, Baifeng Shi, Qinghao Hu et al.NeurIPS 2025 · 91 citations
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun et al.CVPR 2024 · 83 citations
- Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal EvidenceJiahao Meng, Xiangtai Li, Haochen Wang, Tan Yue et al.ICML 2026 · 43 citations
- ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language ModelsIlker Kesen, Andrea Pedrotti, Mustafa Dogan, Michele Cafagna et al.ICLR 2024 · 25 citations
Related papers
- MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering BenchmarkShaden Shaar, Bradon Thymes, Sirawut Chaixanien, Claire Cardie et al.CVPR 2026 · 4 citations
- VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in VideoHanoona Abdul Rasheed, Abdelrahman M Shaker, Anqi Tang, Muhammad Maaz et al.ICLR 2026 · 16 citations
- NExT-QA: Next Phase of Question-Answering to Explaining Temporal ActionsJunbin Xiao, Xindi Shang, Angela Yao, Tat-Seng ChuaCVPR 2021
- MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in VideosKejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li et al.ICLR 2026 · 22 citations
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMsCaorui Li, Yu Chen, Yiyan Ji, Jin Xu et al.ICLR 2026 · 53 citations
