VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR
Shenghui Chen, Po-han Li, Sandeep Chinchali, Ufuk Topcu
摘要
Many decision-making tasks, where both accuracy and efficiency matter, still require human supervision. For example, tasks like traffic officers reviewing hour-long dashcam footage or researchers screening conference videos can benefit from concise summaries that reduce cognitive load and save time. Yet current vision-language models (VLMs) often produce verbose, redundant outputs that hinder task performance. Existing video caption evaluation depends on costly human annotations and overlooks the summaries' utility in downstream tasks. We address these gaps with Video-to-text Information Bottleneck Evaluation (VIBE), an annotation-free method that scores VLM outputs using two metrics: grounding (how well the summary aligns with visual content) and utility (how informative it is for the task). VIBE selects from randomly sampled VLM outputs by ranking them according to the two scores to support effective human decision-making. Human studies on LearningPaper24, SUTD-TrafficQA, and LongVideoBench show that summaries selected by VIBE consistently improve performance-boosting task accuracy by up to 61.23% and reducing response time by 75.77% compared to naive VLM summaries or raw video. 2
- Equal contribution (Order determined by coin toss). 2 Project Website, Code, and LearningPaper24 Dataset. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Video (𝑽) VLM Captioning Task (𝒀) Summaries Candidates (𝑻) "Annotation-Free Rejection Sampling" "Raw Video" This video teaches the physics principle regarding birds conserving body heat due to blood vessels … It references a webpage called "Physics of the Everyday" for learning physics of everyday activities.
The video explores the unique physical adaptations of flamingos, focusing on their ability to stand on one leg and conserve body heat … It features visuals of flamingos throughout the video.
This video from SciShow covers the unique one-legged sleeping stance of flamingos and compares it to a pigeon's roosting behavior, using diagrams and footage to explain their anatomy. It concludes with a promotion for Brilliant's Physics courses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
- How Does Information Bottleneck Help Deep Learning?Kenji Kawaguchi, Zhun Deng, Xu Ji, Jiaoyang HuangICML 2023 · 被引用 117 次
- Video Question Answering: Datasets, Algorithms and ChallengesYaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li 等EMNLP 2022 · 被引用 70 次
- RouteLLM: Learning to Route LLMs from Preference DataIsaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang 等ICLR 2025
相关 Paper
- Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question AnsweringZizhen Wang, Bo Feng, Zhengfeng Lai, Shiyu Li 等ACL 2026
- Explaining A Black-box By Using A Deep Variational Information Bottleneck ApproachSeo-Jin Bang, Pengtao Xie, Heewook Lee, Wei Wu 等AAAI 2021 · 被引用 33 次
- DeVAn: Dense Video Annotation for Video-Language ModelsTingkai Liu, Yunzhe Tao, Haogeng Liu, Qihang Fan 等ACL 2024 · 被引用 1 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language ModelsOlga Loginova, Oleksandr Bezrukov, Ravi Shekhar, Alexey KravetsACL 2025 · 被引用 8 次
