Augmented Partial Mutual Learning with Frame Masking for Video Captioning
Ke Lin, Zhuoxin Gan, Liwei Wang
摘要
Recent video captioning work improves greatly due to the invention of various elaborate model architectures. If multiple captioning models are combined into a unified framework not only by simple more ensemble, and each model can benefit from each other, the final captioning might be boosted further. Jointly training of multiple model have not been explored in previous works. In this paper, we propose a novel Augmented Partial Mutual Learning (APML) training method where multiple decoders are trained jointly with mimicry losses between different decoders and different input variations. Another problem of training captioning model is the "one-to-many" mapping problem which means that one identical video input is mapped to multiple caption annotations. To address this problem, we propose an annotation-wise frame masking approach to convert the "one-to-many" mapping to "one-to-one" mapping. The experiments performed on MSR-VTT and MSVD datasets demonstrate our proposed algorithm achieves the state-of-the-art performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Accommodating Audio Modality in CLIP for Multimodal ProcessingLudan Ruan, Anwen Hu, Yuqing Song, Liang Zhang 等AAAI 2023 · 被引用 18 次
- Token Mixing: Parameter-Efficient Transfer Learning from Image-Language to Video-LanguageYuqi Liu, Luhui Xu, Pengfei Xiong, Qin JinAAAI 2023 · 被引用 10 次
- STOA-VLP: Spatial-Temporal Modeling of Object and Action for Video-Language Pre-trainingWeihong Zhong, Mao Zheng, Duyu Tang, Xuan Luo 等AAAI 2023 · 被引用 9 次
- Self-Critical Distillation Network for Video-based Commonsense CaptioningMengqi Yuan, Gengyun Jia, Bing-Kun BaoCVPR 2026
它引用的顶会 Paper4
- Controllable Video Captioning With POS Sequence Guidance Based on Gated Fusion NetworkBairui Wang, Lin Ma, Wei Zhang, Wenhao Jiang 等ICCV 2019 · 被引用 183 次
- Joint Syntax Representation Learning and Visual Cue Translation for Video CaptioningJingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo 等ICCV 2019 · 被引用 84 次
- Spatio-Temporal Graph for Video Captioning With Knowledge DistillationBoxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee 等CVPR 2020
- Object Relational Graph With Teacher-Recommended Learning for Video CaptioningZiqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li 等CVPR 2020
相关 Paper
- End-to-end Generative Pretraining for Multimodal Video CaptioningPaul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, Cordelia SchmidCVPR 2022 · 被引用 152 次
- MAMS: Model-Agnostic Module Selection Framework for Video CaptioningSangho Lee, Il Yong Chun, Hogun ParkAAAI 2025 · 被引用 1 次
- Text with Knowledge Graph Augmented Transformer for Video CaptioningXin Gu, Guang Chen, Yufei Wang, Libo Zhang 等CVPR 2023
- Semantic Grouping Network for Video CaptioningHobin Ryu, Sunghun Kang, Haeyong Kang, Chang D. YooAAAI 2021 · 被引用 160 次
- Learnability Matters: Active Learning for Video CaptioningYiqian Zhang, Buyu Liu, Jun Bao, Qiang Huang 等NeurIPS 2024 · 被引用 47 次
