VideoExplorer: Advancing Long-Horizon Video Understanding via Hierarchical Orchestration
Huaying Yuan, Zheng Liu, Junjie Zhou, Hongjin Qian, Yan Shu, Nicu Sebe, Ji-Rong Wen, Zhicheng Dou
Abstract
Current agentic frameworks for Long-Video Understanding (LVU) remain limited by two critical problems: ineffective control, where traditional monolithic agents struggle with high-branching, multi-granularity decision processes; and inefficient supervision, where sparse, outcome-based feedback fails to guide long-horizon reasoning. To resolve these challenges, we propose VideoExplorer, a novel agentic system designed to advance long-video reasoning on top of structured control and trajectory-level optimization. First, VideoExplorer innovates a hierarchically orchestrated framework: it employs a planning agent to focus on creating high-level reasoning strategies and specialized sub-agents (including a temporal grounder and a visual perceiver) to accomplish fine-grained reasoning executions, thereby substantially reducing the complexity of reasoning process. Second, VideoExplorer introduces a novel optimization approach, Trajectory level Direct Preference Optimization (TDPO), to mitigate inefficient supervision. Unlike standard methods that optimize single turn responses, TDPO aligns the entire planning trajectory, including evidence routing and termination decisions, with end task success, which effectively mitigates premature commitment and compounding errors. To better support the conduct of TDPO, we further create fine-grained supervision data via a teacher-guided, difficulty-adaptive sampling process. Extensive experiments on MLVU, LVBench, and MH-NIAH demonstrate that VideoExplorer consistently outperforms monolithic baselines in both accuracy and efficiency, validating the effectiveness of structured control and trajectory-level optimization in long-video reasoning. Our code is available in this repository.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9e9b6f34-8e30-49ef-b5df-384e63d5c5c4Related papers
- VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video ReasoningYang Ding, Xin Lai, Yizhen Zhang, Wei Li et al.ICLR 2026 · 26 citations
- VCA: Video Curious Agent for Long Video UnderstandingZeyuan Yang, Delin Chen, Xueyang Yu, Maohao Shen et al.ICCV 2025 · 5 citations
- ReAgent-V: A Reward-Driven Multi-Agent Framework for Video UnderstandingYiyang Zhou, Yangfan He, Yaofeng Su, Siwei Han et al.NeurIPS 2025 · 55 citations
- VideoSeek: Long-Horizon Video Agent with Tool-Guided SeekingJingyang Lin, Jialian Wu, Jiang Liu, Ximeng Sun et al.CVPR 2026 · 15 citations
- Native Active Perception as Reasoning for Omni-Modal UnderstandingZhenghao Xing, Ruiyang Xu, Yuxuan Wang, Jinzheng He et al.ICML 2026 · 1 citation
