Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
Xiaoyi Zhang, Zhaoyang Jia, Zongyu Guo, Jiahao Li, Bin Li, Houqiang Li, Yan Lu
摘要
Long-form video understanding presents significant challenges due to extensive temporal-spatial complexity and the difficulty of question answering under such extended contexts. While Large Language Models (LLMs) have demonstrated considerable advancements in video analysis capabilities and long context handling, they continue to exhibit limitations when processing information-dense hour-long videos. To overcome such limitations, we propose the Deep Video Discovery (DVD) agent to leverage an agentic search strategy over segmented video clips. Unlike previous video agents that rely on predefined workflows applied uniformly across different queries, our approach emphasizes the autonomous and adaptive nature of agents. By providing a set of search-centric tools on multi-granular video database, our DVD agent leverages the advanced reasoning capability of LLM to plan on its current observation state, strategically selects tools to orchestrate adaptive workflow for different queries in light of the gathered information. We perform comprehensive evaluation on multiple long video understanding benchmarks that demonstrates our advantage. Our DVD agent achieves state-of-the-art performance on the challenging LVBench dataset, reaching an accuracy of 74.2%, which substantially surpasses all prior works, and further improves to 76.0% with transcripts. The code has been released at https://github.com/microsoft/DeepVideoDiscovery. * Equal contribution. † This work was done during the internship at Microsoft Research Asia as an open-source project. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame SpotlightingZefeng He, Xiaoye Qu, Yafu Li, Siyuan Huang 等ICLR 2026 · 被引用 34 次
- VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video ReasoningYang Ding, Xin Lai, Yizhen Zhang, Wei Li 等ICLR 2026 · 被引用 26 次
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data SynthesisCongzhi Zhang, Zhibin Wang, Yinchao Ma, Jiawei Peng 等ICLR 2026 · 被引用 24 次
- VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video UnderstandingYufei Yin, Qianke Meng, Minghao Chen, Jiajun Ding 等CVPR 2026 · 被引用 24 次
- VideoSeek: Long-Horizon Video Agent with Tool-Guided SeekingJingyang Lin, Jialian Wu, Jiang Liu, Ximeng Sun 等CVPR 2026 · 被引用 15 次
它引用的顶会 Paper12
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- VideoChat-Flash: Hierarchical Compression for Long-Context Video ModelingXinhao Li, Yi Wang, Jiashuo Yu, Xiangyu Zeng 等ICLR 2026 · 被引用 172 次
- LVBench: An Extreme Long Video Understanding BenchmarkWeihan Wang, Zehai He, Wenyi Hong, Yean Cheng 等ICCV 2025 · 被引用 28 次
- VCA: Video Curious Agent for Long Video UnderstandingZeyuan Yang, Delin Chen, Xueyang Yu, Maohao Shen 等ICCV 2025 · 被引用 5 次
- AVA: Towards Agentic Video Analytics with Vision Language ModelsYuxuan Yan, Shiqi Jiang, Ting Cao, Yifan Yang 等NSDI 2026 · 被引用 4 次
相关 Paper
- LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM AgentsBoyu Chen, Zhengrong Yue, Siran Chen, Zikang Wang 等ICCV 2025 · 被引用 12 次
- MR. Video: MapReduce as an Effective Principle for Long Video UnderstandingZiqi Pang, Yu-Xiong WangNeurIPS 2025 · 被引用 6 次
- LensWalk: Agentic Video Understanding by Planning How You See in VideosKeliang Li, Yansong Li, Hongze Shen, Mengdi Liu 等CVPR 2026 · 被引用 13 次
- DrVideo: Document Retrieval Based Long Video UnderstandingZiyu Ma, Chenhui Gou, Hengcan Shi, Bin Sun 等CVPR 2025
- LongVideo-R1: Smart Navigation for Low-cost Long Video UnderstandingJihao Qiu, Lingxi Xie, Xinyue Huo, Qi Tian 等CVPR 2026 · 被引用 7 次
