NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic Reasoning
Sahil Shah, S. P. Sharan, Harsh Goel, Minkyu Choi, Mustafa Munir, Manvik Pasula, Radu Marculescu, Sandeep Chinchali
Abstract
While vision-language models (VLMs) excel at tasks involving single images or short videos, they still struggle with Long Video Question Answering (LVQA) due to its demand for complex multi-step temporal reasoning. Vanilla approaches, which simply sample frames uniformly and feed them to a VLM along with the question, incur significant token overhead. This forces aggressive downsampling of long videos, causing models to miss fine-grained visual structure, subtle event transitions, and key temporal cues. Recent works attempt to overcome these limitations through heuristic approaches; however, they lack explicit mechanisms for encoding temporal relationships and fail to provide any formal guarantees that the sampled context actually encodes the compositional or causal logic required by the question. To address these foundational gaps, we introduce NeuS-QA, a training-free, plug-and-play neuro-symbolic pipeline for LVQA. NeuS-QA first translates a natural language question into a logic specification that models the temporal relationship between frame-level events. Next, we construct a video automaton to model the video's frame-by-frame event progression, and finally employ model checking to compare the automaton against the specification to identify all video segments that satisfy the question's logical requirements. Only these logic-verified segments are submitted to the VLM, thus improving interpretability, reducing hallucinations, and enabling compositional reasoning without modifying or fine-tuning the model. Experiments on the LongVideoBench and CinePile LVQA benchmarks show that NeuS-QA significantly improves performance by over 10%, particularly on questions involving event ordering, causality, and multi-step reasoning. We open-source our code at https://utaustin-swarmlab.github.io/NeuS-QA/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bae00ebf-8154-46a3-bccf-8d8b17270ad3Builds on7
- A Graph-Based Framework to Bridge Movies and SynopsesYu Xiong, Qingqiu Huang, Lingfeng Guo, Hang Zhou et al.ICCV 2019 · 71 citations
- LVBench: An Extreme Long Video Understanding BenchmarkWeihan Wang, Zehai He, Wenyi Hong, Yean Cheng et al.ICCV 2025 · 28 citations
- ALLVB: All-in-One Long Video Understanding BenchmarkXichen Tan, Yuanjing Luo, Yunfan Ye, Fang Liu et al.AAAI 2025 · 13 citations
- BIMBA: Selective-Scan Compression for Long-Range Video Question AnsweringMd Mohaiminul Islam, Tushar Nagarajan, Huiyu Wang, Gedas Bertasius et al.CVPR 2025
- Neuro-Symbolic Evaluation of Text-to-Video Models using Formal VerificationS. P. Sharan, Minkyu Choi, Sahil Shah, Harsh Goel et al.CVPR 2025
Related papers
- M-LLM Based Video Frame Selection for Efficient Video UnderstandingKai Hu, Feng Gao, Xiaohan Nie, Peng Zhou et al.CVPR 2025
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMsMinji Kim, Taekyung Kim, Bohyung HanICLR 2026 · 8 citations
- STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-TrainingHaiyi Qiu, Minghe Gao, Long Qian, Kaihang Pan et al.CVPR 2025
- A Training-Free Framework for Long Video Understanding via Video-Query-Options SimilarityZhirong Wu, Xiaodong Wang, Langling Huang, Teng Xu et al.ICLR 2026
- VTimeCoT: Thinking by Drawing for Video Temporal Grounding and ReasoningJinglei Zhang, Yuanfan Guo, Rolandos Alexandros Potamias, Jiankang Deng et al.ICCV 2025 · 4 citations
