VITED: Video Temporal Evidence Distillation
Yujie Lu, Yale Song, William Wang, Lorenzo Torresani, Tushar Nagarajan
Abstract
We investigate complex video question answering via chainof-evidence reasoning -identifying sequences of temporal spans from multiple relevant parts of the video, together with visual evidence within them. Existing models struggle with multi-step reasoning as they uniformly sample a fixed number of frames, which can miss critical evidence distributed nonuniformly throughout the video. Moreover, they lack the ability to temporally localize such evidence in the broader context of the full video, which is required for answering complex questions. We propose a framework to enhance existing VideoQA datasets with evidence reasoning chains, automatically constructed by searching for optimal intervals of interest in the video with supporting evidence, that maximizes the likelihood of answering a given question. We train our model (VITED) to generate these evidence chains directly, enabling it to both localize evidence windows as well as perform multi-step reasoning across them in long-form video content. We show the value of our evidence-distilled models on a suite of long video QA benchmarks where we outperform state-of-the-art approaches that lack evidence reasoning capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2b78fd0-5047-44af-a671-41f64d0b1e0fCited by top-tier papers4
- When Thinking Drifts: Evidential Grounding for Robust Video ReasoningRomy Luo, Zihui Xue, Alex Dimakis, Kristen GraumanNeurIPS 2025 · 21 citations
- Agentic Very Long Video UnderstandingAniket Rege, Arka Sadhu, Yuliang Li, Kejie Li et al.ACL 2026 · 8 citations
- AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision–Language ModelsShih-Po Lee, Ehsan ElhamifarCVPR 2026 · 3 citations
- MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language ModelsYifan Xu, Chao Zhang, Ruifei Ma, Fei Gao et al.CVPR 2026
Builds on27
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
Related papers
- Enhancing Video-LLM Reasoning via Agent-of-Thoughts DistillationYudi Shi, Shangzhe Di, Qirui Chen, Weidi XieCVPR 2025
- GraphVideoAgent: Enhancing Long-form Video Understanding with Entity Relation GraphsMeng Chu, Yicong Li, Tat-Seng ChuaACM MM 2025 · 1 citation
- MIST : Multi-modal Iterative Spatial-Temporal Transformer for Long-form Video Question AnsweringDifei Gao, Luowei Zhou, Lei Ji, Linchao Zhu et al.CVPR 2023
- Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal CluesYan Zhang, Gangyan Zeng, Huawen Shen, Daiqing Wu et al.AAAI 2025 · 1 citation
- Grounded Multi-Hop VideoQA in Long-Form Egocentric VideosQirui Chen, Shangzhe Di, Weidi XieAAAI 2025 · 35 citations
