Self-Critical Distillation Network for Video-based Commonsense Captioning
Mengqi Yuan, Gengyun Jia, Bing-Kun Bao
摘要
Video-based commonsense captioning aims to generate captions for the video content while providing multiple commonsense about the underlying events. Existing approaches rely on constructing a "video → content caption → commonsense" reasoning chain, which generates visually ungrounded commonsense and neglects inter-category commonsense correlations. Firstly, the existing reasoning chain induces the model's excessive reliance on content caption when generating commonsense, resulting in generic outputs with limited visual relevance. Secondly, the reasoning chain adopts multiple isolated decoders for commonsense generation, which fails to leverage the correlations between different categories of commonsense. To address these limitations, we introduce a novel self-critical distillation network (SCD-Net), which optimizes the reasoning chain by enhancing visual reasoning and establishing inter-category commonsense correlations. Specifically, on the one hand, we introduce self-critical learning and design a reward function to refine model outputs, encouraging full use of visual information and enhancing visual understanding. On the other hand, we propose a joint reasoning distillation framework that fosters mutual inference among diverse commonsense categories. In this framework, we incorporate the cascaded decoder and knowledge distillation strategy to facilitate inter-category commonsense knowledge transfer while maintaining the fairness of the testing. Our experiments on the large-scale Video-to-Commonsense dataset demonstrate that our approach performs favorably against stateof-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsKunchang Li, Yali Wang, Yizhuo Li, Yi Wang 等ICCV 2023 · 被引用 266 次
- End-to-end Generative Pretraining for Multimodal Video CaptioningPaul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, Cordelia SchmidCVPR 2022 · 被引用 152 次
- Video2Commonsense: Generating Commonsense Descriptions to Enrich Video CaptioningZhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral 等EMNLP 2020 · 被引用 61 次
- Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual RecognitionChuanguang Yang, Xinqiang Yu, Han Yang, Zhulin An 等AAAI 2025 · 被引用 26 次
- Augmented Partial Mutual Learning with Frame Masking for Video CaptioningKe Lin, Zhuoxin Gan, Liwei WangAAAI 2021 · 被引用 25 次
相关 Paper
- Hybrid Reasoning Network for Video-based Commonsense CaptioningWeijiang Yu, Jian Liang, Lei Ji, Lu Li 等ACM MM 2021 · 被引用 8 次
- Multi-Domain Lifelong Visual Question Answering via Self-Critical DistillationMingrui Lao, Nan Pu, Yu Liu, Zhun Zhong 等ACM MM 2023 · 被引用 5 次
- Video Event Extraction with Multi-View Interaction Knowledge DistillationKaiwen Wei, Runyan Du, Li Jin, Jian Liu 等AAAI 2024 · 被引用 5 次
- Joint Commonsense and Relation Reasoning for Image and Video CaptioningJingyi Hou, Xinxiao Wu, Xiaoxun Zhang, Yayun Qi 等AAAI 2020 · 被引用 52 次
- From Representation to Reasoning: Towards both Evidence and Commonsense Reasoning for Video Question-AnsweringJiangtong Li, Li Niu, Liqing ZhangCVPR 2022 · 被引用 48 次
