Hierarchical Multi-Source Uncertainty Aggregation for Interactive Video Captioning
Ervine Zheng, Qi Yu
摘要
Video captioning automatically generates natural language phrases to explain the contents in video frames. When deploying captioning models in specialized domains, active learning can help reduce the high annotation cost. However, the generative nature of the captioning process is more complex than standard supervised learning tasks and introduces several challenges for active learning in video captioning. Entropy-based uncertainty estimation, which is widely used in active learning, may be inflated in captioning tasks and mislead active sampling. Another challenge arises from the rich content of videos, as each video could be described in multiple ways. A single uncertainty score obtained from one possible caption does not capture the diversity induced by the rich content. To fill out this gap, we propose identifying multiple sources of uncertainty and performing hierarchical aggregation to integrate uncertainty from distinct sources. This innovates a holistic uncertainty metric to quantify the overall informativeness of video content for active sampling. The overall uncertainty is built upon conditional vacuity, an extension of the second-order uncertainty introduced along with the evidential learning framework to the captioning setting, leading to more robust uncertainty estimation without inflation. Both theoretical analysis and experimental evaluation are conducted to demonstrate the effectiveness of the proposed framework for complex uncertainty estimation and interactive learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Deep Evidential RegressionAlexander Amini, Wilko Schwarting, Ava Soleimany, Daniela RusNeurIPS 2020 · 被引用 777 次
相关 Paper
- Learnability Matters: Active Learning for Video CaptioningYiqian Zhang, Buyu Liu, Jun Bao, Qiang Huang 等NeurIPS 2024 · 被引用 47 次
- Are Binary Annotations Sufficient? Video Moment Retrieval via Hierarchical Uncertainty-based Active LearningWei Ji, Renjie Liang, Zhedong Zheng, Wenqiao Zhang 等CVPR 2023
- Multifaceted Uncertainty Estimation for Label-Efficient Deep LearningWeishi Shi, Xujiang Zhao, Feng Chen, Qi YuNeurIPS 2020 · 被引用 38 次
- Evidential Interactive Learning for Medical Image CaptioningErvine Zheng, Qi YuICML 2023 · 被引用 10 次
- Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval Via Uncertainty MinimizationBingqing Zhang, Zhuo Cao, Heming Du, Yang Li 等ICCV 2025 · 被引用 3 次
