VideoExplorer: Advancing Long-Horizon Video Understanding via Hierarchical Orchestration
Huaying Yuan, Zheng Liu, Junjie Zhou, Hongjin Qian, Yan Shu, Nicu Sebe, Ji-Rong Wen, Zhicheng Dou
摘要
Current agentic frameworks for Long-Video Understanding (LVU) remain limited by two critical problems: ineffective control, where traditional monolithic agents struggle with high-branching, multi-granularity decision processes; and inefficient supervision, where sparse, outcome-based feedback fails to guide long-horizon reasoning. To resolve these challenges, we propose VideoExplorer, a novel agentic system designed to advance long-video reasoning on top of structured control and trajectory-level optimization. First, VideoExplorer innovates a hierarchically orchestrated framework: it employs a planning agent to focus on creating high-level reasoning strategies and specialized sub-agents (including a temporal grounder and a visual perceiver) to accomplish fine-grained reasoning executions, thereby substantially reducing the complexity of reasoning process. Second, VideoExplorer introduces a novel optimization approach, Trajectory level Direct Preference Optimization (TDPO), to mitigate inefficient supervision. Unlike standard methods that optimize single turn responses, TDPO aligns the entire planning trajectory, including evidence routing and termination decisions, with end task success, which effectively mitigates premature commitment and compounding errors. To better support the conduct of TDPO, we further create fine-grained supervision data via a teacher-guided, difficulty-adaptive sampling process. Extensive experiments on MLVU, LVBench, and MH-NIAH demonstrate that VideoExplorer consistently outperforms monolithic baselines in both accuracy and efficiency, validating the effectiveness of structured control and trajectory-level optimization in long-video reasoning. Our code is available in this repository.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video ReasoningYang Ding, Xin Lai, Yizhen Zhang, Wei Li 等ICLR 2026 · 被引用 26 次
- VCA: Video Curious Agent for Long Video UnderstandingZeyuan Yang, Delin Chen, Xueyang Yu, Maohao Shen 等ICCV 2025 · 被引用 5 次
- ReAgent-V: A Reward-Driven Multi-Agent Framework for Video UnderstandingYiyang Zhou, Yangfan He, Yaofeng Su, Siwei Han 等NeurIPS 2025 · 被引用 55 次
- VideoSeek: Long-Horizon Video Agent with Tool-Guided SeekingJingyang Lin, Jialian Wu, Jiang Liu, Ximeng Sun 等CVPR 2026 · 被引用 15 次
- Native Active Perception as Reasoning for Omni-Modal UnderstandingZhenghao Xing, Ruiyang Xu, Yuxuan Wang, Jinzheng He 等ICML 2026 · 被引用 1 次
