Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-Learning
Zhuyang Xie, Yan Yang, Yankai Yu, Jie Wang, Yongquan Jiang, Xiao Wu
Abstract
Dense video captioning aims to detect and describe all events in untrimmed videos. This paper presents a dense video captioning network called Multi-Concept Cyclic Learning (MCCL), which aims to: (1) detect multiple concepts at the frame level and leverage these concepts to provide temporal event cues; and (2) establish cyclic co-learning between the generator and the localizer within the captioning network to promote semantic perception and event localization. Specifically, weakly supervised concept detection is performed for each frame, and the detected concept embeddings are integrated into the video features to provide event cues. Additionally, video-level concept contrastive learning is introduced to produce more discriminative concept embeddings. In the captioning network, a cyclic co-learning strategy is proposed, where the generator guides the localizer for event localization through semantic matching, while the localizer enhances the generator’s event semantic perception through location matching, making semantic perception and event localization mutually beneficial. MCCL achieves state-of-the-art performance on the ActivityNet Captions and YouCook2 datasets. Extensive experiments demonstrate its effectiveness and interpretability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a447003-0c66-44c7-b76f-83f99b9dba6cCited by top-tier papers2
- Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video CaptioningSeungHyup Baek, Jimin Lee, Hyeongkeun Lee, Jae Won ChoCVPR 2026 · 1 citation
- Task-Specific Information Decomposition for End-to-End Dense Video CaptioningZhiyue Liu, Xinru Zhang, Jinyuan LiuACL 2025
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng et al.ICCV 2021 · 238 citations
- Semantic Grouping Network for Video CaptioningHobin Ryu, Sunghun Kang, Haeyong Kang, Chang D. YooAAAI 2021 · 160 citations
- Watch, Listen and Tell: Multi-Modal Weakly Supervised Dense Event CaptioningTanzila Rahman, Bicheng Xu, Leonid SigalICCV 2019 · 89 citations
- Viewpoint-Agnostic Change Captioning with Cycle ConsistencyHoeseong Kim, Jongseok Kim, Hyungseok Lee, Hyunsung Park et al.ICCV 2021 · 56 citations
Related papers
- Towards Bridging Event Captioner and Sentence Localizer for Weakly Supervised Dense Event CaptioningShaoxiang Chen, Yu-Gang JiangCVPR 2021
- Hierarchical Context-aware Network for Dense Video Event CaptioningLei Ji, Xianglin Guo, Haoyang Huang, Xilin ChenACL 2021
- Do You Remember? Dense Video Captioning with Cross-Modal Memory RetrievalMinkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi et al.CVPR 2024
- Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal InteractionMingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li et al.AAAI 2026 · 1 citation
- Exploiting Auxiliary Caption for Video GroundingHongxiang Li, Meng Cao, Xuxin Cheng, Yaowei Li et al.AAAI 2024 · 16 citations
