Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering
Yumeng Shi, Quanyu Long, Wenya Wang
Abstract
Video question answering benefits from the rich information in videos, enabling various applications. However, the large volume of tokens generated from long videos presents challenges to memory efficiency and model performance. To alleviate this, existing works propose to compress video inputs, but often overlook the varying importance of static and dynamic information across different queries, leading to inefficient token usage within limited budgets. We propose a novel token selection strategy, EXPLORE-THEN-SELECT, that adaptively adjusts static and dynamic information based on question requirements. Our framework first explores different token allocations between key frames, which preserve spatial details, and delta frames, which capture temporal changes. Then it employs a query-aware attention-based metric to select the optimal token combination without model updates. Our framework is plug-and-play and can be seamlessly integrated within diverse video language models. Extensive experiments show that our method achieves significant performance improvements (up to 5.8%) on multiple video question answering benchmarks. Our code is available at https://github.com/ANDgate99/Explore-Then-Select .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02801151-60b0-45f9-9beb-2f86e6843a6aCited by top-tier papers5
- Causality Matters: How Temporal Information Emerges in Video Language ModelsYumeng Shi, Quanyu Long, Yin Wu, Wenya WangAAAI 2026 · 4 citations
- LLMC+: Benchmarking Vision-Language Model Compression with a plug-and-play ToolkitChengtao Lv, Bilang Zhang, Yang Yong, Ruihao Gong et al.AAAI 2026 · 3 citations
- CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMsJingyu Lei, Gaoang Wang, Der-Horng LeeCVPR 2026 · 1 citation
- ReGATE: Learning Faster and Better with Fewer Tokens in MLLMsChaoyu Li, Yogesh Kulkarni, Pooyan FazliACL 2026
- APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate AttentionYuxiang Huang, Mingye Li, Xu Han, Chaojun Xiao et al.ACL 2026
Builds on9
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Video-LLaVA: Learning United Visual Representation by Alignment Before ProjectionBin Lin, Yang Ye, Bin Zhu, Jiaxi Cui et al.EMNLP 2024 · 231 citations
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.NeurIPS 2024 · 216 citations
- MovieChat: From Dense Token to Sparse Memory for Long Video UnderstandingEnxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang et al.CVPR 2024 · 95 citations
- Token Merging: Your ViT But FasterDaniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang et al.ICLR 2023 · 62 citations
Related papers
- Unified Spatiotemporal Token Compression for Video-LLMs at Ultra-Low RetentionJunhao Du, Jialong Xue, Anqi Li, Jincheng Dai et al.CVPR 2026 · 7 citations
- Granularity-Adaptive Spatial Evidence Tokenization for Video Question AnsweringHao Jiang, Yang Jin, Zhicheng Sun, Kun Xu et al.AAAI 2025 · 2 citations
- FlexSelect: Flexible Token Selection for Efficient Long Video UnderstandingYunzhu Zhang, Yu Lu, Tianyi Wang, Fengyun Rao et al.NeurIPS 2025 · 22 citations
- Vista-LLM: Decoupled Query-Guided Visual Token Pruning for Efficient Long-Video Large Language ModelsZhenyu Li, Zuchao Li, Ping Wang, Lefei Zhang et al.ACL 2026
- QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video ComprehensionYongdong Luo, Wang Chen, Weizhong Huang, Shukang Yin et al.AAAI 2026
