VAST: Video Ability-Stratified Taxonomy for Data-Efficient Video Reasoning
Zhongan Wang, Xiaoyu Wen, Lingxiao Du, Kun Li, Zhiliang Wu, Xingcheng Xu, Qiaosheng Zhang, Chaochao Lu, Hehe Fan
摘要
Reinforcement learning (RL) has emerged as an effective approach for improving video reasoning in multimodal large language models (MLLMs). However, existing methods remain inefficient for two reasons. First, training data are typically organized by task formats rather than underlying reasoning abilities, creating a mismatch that encourages models to learn task-specific patterns instead of transferable capabilities. As a result, improving ability generalization requires broad coverage over many ability-task combinations, making RL training costly. Second, prior work often compensates for this inefficiency with increasingly complex algorithmic designs, such as specialized temporal architectures or multi-objective reward frameworks, which further complicate training. To address these issues, we present VAST, a cognitive taxonomy and data organization framework that structures video understanding into three layers: Perception, Reasoning, and Cognition. Based on this taxonomy, we construct VAST-15K for training and VAST-Bench for evaluation. We further introduce Video-VAST, a reinforcement learning framework that uses consistency rewards to encourage alignment between reasoning traces and final answers, without architectural modifications. Experiments show that VideoVAST achieves 66.3% accuracy on MVBench and 57.4% on VAST-Bench, compared with 62.7% and 54.3% for Video-R1, while using approximately 72% fewer GPU hours and 96% fewer training samples under the same training settings. Code and resources are available at https://zhongan-wang. github.io/VAST.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
相关 Paper
- Incentivizing Versatile Video Reasoning in MLLMs via Data-Efficient Reinforcement LearningXiaodong Wang, Zhirong Wu, Langling Huang, Yuxi Zheng 等CVPR 2026
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video ReasoningHaoji Zhang, Xin Gu, Jiawen Li, Chixiang Ma 等CVPR 2026 · 被引用 92 次
- Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual EvidenceKun Ouyang, Yuanxin Liu, Linli Yao, Yishuo Cai 等CVPR 2026 · 被引用 17 次
- Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-AnsweringZhaohe Liao, Jiangtong Li, Siyu Sun, Qingyang Liu 等ICML 2025
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data SynthesisCongzhi Zhang, Zhibin Wang, Yinchao Ma, Jiawei Peng 等ICLR 2026 · 被引用 24 次
