VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
Ziyang Wang, Shoubin Yu, Elias Stengel-Eskin, Jaehong Yoon, Feng Cheng, Gedas Bertasius, Mohit Bansal
Abstract
Long-form video understanding is complicated by the high redundancy of video data and the abundance of queryirrelevant information. To tackle these challenges, we propose VIDEOTREE, a training-free framework which builds a query-adaptive and hierarchical video representation for LLM reasoning over long-form videos. First, VIDEOTREE extracts query-relevant information from the input video through an iterative process, progressively refining the selection of keyframes based on their relevance to the query. Furthermore, VIDEOTREE leverages the inherent hierarchical structure of long video data, which is often overlooked by existing LLM-based methods. Specifically, we incorporate multi-granularity information into a tree-based representation, allowing VIDEOTREE to extract query-relevant details from long videos in a coarse-to-fine manner. This enables the model to effectively handle a wide range of video queries with varying levels of detail. Finally, VIDEOTREE aggregates the hierarchical query-relevant information within the tree structure and feeds it into an LLM reasoning model to answer the query. Our experiments show that our method improves both reasoning accuracy and efficiency. Specifically, VIDEOTREE outperforms existing training-free approaches on EgoSchema and NExT-QA with less inference time, achieving 61.1% and 75.6% accuracy on the test set without additional video-specific training. Moreover, on the long split of Video-MME (average 44 minutes), VIDEOTREE achieves better performance than GPT-4V and many other MLLMs that were extensively trained on video data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ddbb478f-f092-4611-b465-04bcbb00db4aCited by top-tier papers78
- Deep Video Discovery: Agentic Search with Tool Use for Long-form Video UnderstandingXiaoyi Zhang, Zhaoyang Jia, Zongyu Guo, Jiahao Li et al.NeurIPS 2025 · 95 citations
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative PerceptionZiang Yan, Yinan He, Xinhao Li, Zhengrong Yue et al.NeurIPS 2025 · 70 citations
- WorldMM: Dynamic Multimodal Memory Agent for Long Video ReasoningWoongyeong Yeo, Kangsan Kim, Jaehong Yoon, Sung Ju HwangCVPR 2026 · 53 citations
- Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video UnderstandingXiaoqian Shen, Wenxuan Zhang, Jun Chen, Mohamed ElhoseinyNeurIPS 2025 · 37 citations
- Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video UnderstandingWeiyu Guo, Ziyang Chen, Shaoguang Wang, JianXiang He et al.NeurIPS 2025 · 35 citations
Builds on34
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 732 citations
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan et al.EMNLP 2020 · 387 citations
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 281 citations
Related papers
- M-LLM Based Video Frame Selection for Efficient Video UnderstandingKai Hu, Feng Gao, Xiaohan Nie, Peng Zhou et al.CVPR 2025
- A Training-Free Framework for Long Video Understanding via Video-Query-Options SimilarityZhirong Wu, Xiaodong Wang, Langling Huang, Teng Xu et al.ICLR 2026
- Q-Frame: Query-Aware Frame Selection and Multi-Resolution Adaptation for Video-LLMsShaojie Zhang, Jiahui Yang, Jianqin Yin, Zhenbo Luo et al.ICCV 2025 · 15 citations
- Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-AnsweringZhaohe Liao, Jiangtong Li, Siyu Sun, Qingyang Liu et al.ICML 2025
- KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMsBaiyang Song, Jun Peng, Yuxin Zhang, Guangyao Chen et al.AAAI 2026
