VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
Chenhao Qiu, Yechao Zhang, Xin Luo, Shien Song, Xusheng Liu
Abstract
Long video question answering requires locating sparse, time-scattered visual evidence within highly redundant content. Although current MLLMs perform well on short videos, long videos introduce long-horizon search and verification, which often necessitates multi-turn, agentic interaction. We show that existing LVU agents can exhibit evidence misalignment: they produce correct answers that are not supported by the retrieved or inspected evidence. To characterize this failure, we introduce two diagnostics (temporal groundedness and semantic groundedness) and use them to reveal two pressures that amplify misalignment: prompt pressure from shared-context saturation at inference time and reward pressure from outcome-only optimization during training. These findings point to a structural root cause: the coupled agent paradigm conflates long-horizon planning with answer authority. We therefore propose the decoupled planner-inspector framework, which separates planning from answer authority and gates final answering on pixel-level verification. Across four long-video benchmarks, our framework improves both answer accuracy and evidence alignment, achieving 55.1% on LVBench and 62.0% on LongVideoBench while producing interpretable search trajectories. Moreover, the decoupled architecture scales consistently with increased search budgets and supports plug-andplay upgrades of the MLLM backbone without retraining the planner. Code and models are available at https://github.com/Echochef/ VideoSEAL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a542e3a-1f88-4bdd-9b9f-e387157f53a8Builds on20
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo et al.NeurIPS 2025 · 528 citations
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 281 citations
- Deep Video Discovery: Agentic Search with Tool Use for Long-form Video UnderstandingXiaoyi Zhang, Zhaoyang Jia, Zongyu Guo, Jiahao Li et al.NeurIPS 2025 · 95 citations
Related papers
- Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop ReasoningXiangyu Zeng, Zhiqiu Zhang, Yuhan Zhu, Xinhao Li et al.ICML 2026 · 14 citations
- LongVideoAgent: Multi-Agent Reasoning with Long VideosRuntao Liu, Ziyi Liu, Jiaqi Tang, Yue Ma et al.ACL 2026 · 17 citations
- Think, Then Verify: A Hypothesis-Verification Multi-Agent Framework for Long Video UnderstandingZheng Wang, Haoran Chen, Haoxuan Qin, Zhipeng Wei et al.CVPR 2026 · 10 citations
- VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video ReasoningYe Liu, Kevin Qinghong Lin, Chang Wen Chen, Mike Zheng ShouICLR 2026 · 23 citations
- Grounded Multi-Hop VideoQA in Long-Form Egocentric VideosQirui Chen, Shangzhe Di, Weidi XieAAAI 2025 · 35 citations
