Towards Sparse Video Understanding and Reasoning
Chenwei Xu, Zhen Ye, Shang Wu, Weijian Li, Zihan Wang, Zhuofan Xia, Lie Lu, Pranav Maneriker, Fan Du, Manling Li, Han Liu
摘要
We present REVISE (Reasoning with Video Sparsity), a multi-round agent for video question answering (VQA). Instead of uniformly sampling frames, REVISE selects a small set of informative frames, maintains a summary-as-state across rounds, and stops early when confident. It supports proprietary vision-language models (VLMs) in a "plugand-play" setting and enables reinforcement fine-tuning for open-source models. For fine-tuning, we introduce EA-GER (Evidence-Adjusted Gain for Efficient Reasoning), an annotation-free reward with three terms: (1) Confidence gain: after new frames are added, we reward the increase in the log-odds margin between the correct option and the strongest alternative; (2) Summary sufficiency: at answer time we re-ask using only the last committed summary and reward success; (3) Correct-and-early stop: answering correctly within a small turn budget is rewarded. Across multiple VQA benchmarks, REVISE improves accuracy while reducing frames, rounds, and prompt tokens, demonstrating practical sparse video reasoning. Project page: https: //sparsevideounderstanding.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper38
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong 等NeurIPS 2024 · 被引用 858 次
相关 Paper
- ReAgent-V: A Reward-Driven Multi-Agent Framework for Video UnderstandingYiyang Zhou, Yangfan He, Yaofeng Su, Siwei Han 等NeurIPS 2025 · 被引用 55 次
- Reinforcing Structured Chain-of-Thought for Video UnderstandingPeiyao Wang, Haotian Xu, Noranart Vesdapunt, Rui Hou 等CVPR 2026 · 被引用 1 次
- From Wrong To Right: A Recursive Approach Towards Vision-Language ExplanationJiaxin Ge, Sanjay Subramanian, Trevor Darrell, Boyi LiEMNLP 2023 · 被引用 4 次
- Select Less, Reason More: Prioritizing Evidence Purity for Video ReasoningXuchen Li, Xuzhao Li, Shiyu Hu, Kaiqi HuangCVPR 2026 · 被引用 5 次
- Reinforce to Learn, Elect to Reason: A Dual Paradigm for Video ReasoningSongyuan Yang, Weijiang Yu, Jilin Ma, Ziyu Liu 等CVPR 2026 · 被引用 1 次
