Towards Diverse Paragraph Captioning for Untrimmed Videos
Yuqing Song, Shizhe Chen, Qin Jin
摘要
Video paragraph captioning aims to describe multiple events in untrimmed videos with descriptive paragraphs. Existing approaches mainly solve the problem in two steps: event detection and then event captioning. Such two-step manner makes the quality of generated paragraphs highly dependent on the accuracy of event proposal detection which is already a challenging task. In this paper, we propose a paragraph captioning model which eschews the problematic event detection stage and directly generates paragraphs for untrimmed videos. To describe coherent and diverse events, we propose to enhance the conventional temporal attention with dynamic video memories, which progressively exposes new video features and suppresses over-accessed video contents to control visual focuses of the model. In addition, a diversity-driven training strategy is proposed to improve diversity of paragraph on the language perspective. Considering that untrimmed videos generally contain massive but redundant frames, we further augment the video encoder with keyframe awareness to improve efficiency. Experimental results on the ActivityNet and Charades datasets show that our proposed model significantly outperforms the state-of-the-art performance on both accuracy and diversity metrics without using any event boundary annotations. Code will be released at https: //github.com/syuqings/video-paragraph . * Equal contribution. This work was performed when Shizhe Chen was at Renmin University of China.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the WildYuanyuan Liu, Wei Dai, Chuanxu Feng, Wenbin Wang 等ACM MM 2022 · 被引用 83 次
- Visual Abductive ReasoningChen Liang, Wenguan Wang, Tianfei Zhou, Yi YangCVPR 2022 · 被引用 50 次
- VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph CaptioningKashu Yamazaki, Khoa Vo, Quang Sang Truong, Bhiksha Raj 等AAAI 2023 · 被引用 44 次
- Learning to Retrieve Videos by Asking QuestionsAvinash Madasu, Junier Oliva, Gedas BertasiusACM MM 2022 · 被引用 17 次
- Movie101: A New Movie Understanding BenchmarkZihao Yue, Qi Zhang, Anwen Hu, Liang Zhang 等ACL 2023 · 被引用 11 次
它引用的顶会 Paper5
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
- Neural Text Generation With Unlikelihood TrainingSean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan 等ICLR 2020 · 被引用 683 次
- MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph CaptioningJie Lei, Liwei Wang, Yelong Shen, Dong Yu 等ACL 2020 · 被引用 168 次
- Spatio-Temporal Graph for Video Captioning With Knowledge DistillationBoxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee 等CVPR 2020
- Object Relational Graph With Teacher-Recommended Learning for Video CaptioningZiqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li 等CVPR 2020
相关 Paper
- Sketch, Ground, and Refine: Top-Down Dense Video CaptioningChaorui Deng, Shizhe Chen, Da Chen, Yuan He 等CVPR 2021
- Hierarchical Context-aware Network for Dense Video Event CaptioningLei Ji, Xianglin Guo, Haoyang Huang, Xilin ChenACL 2021
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng 等ICCV 2021 · 被引用 238 次
- Event-Equalized Dense Video CaptioningKangyi Wu, Pengna Li, Jingwen Fu, Yizhe Li 等CVPR 2025
- Do You Remember? Dense Video Captioning with Cross-Modal Memory RetrievalMinkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi 等CVPR 2024
