DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
Hao Wu, Huabin Liu, Yu Qiao, Xiao Sun
Abstract
We present Dive Into the BoundarieS (DIBS), a novel pretraining framework for dense video captioning (DVC), that elaborates on improving the quality of the generated event captions and their associated pseudo event bound-aries from unlabeled videos. By leveraging the capabil-ities of diverse large language models (LLMs), we gen-erate rich DVC-oriented caption candidates and optimize the corresponding pseudo boundaries under several metic-ulously designed objectives, considering diversity, event-centricity, temporal ordering, and coherence. Moreover, we further introduce a novel online boundary refinement strat-egy that iteratively improves the quality of pseudo bound-aries during training. Comprehensive experiments have been conducted to examine the effectiveness of the pro-posed technique components. By leveraging a substantial amount of unlabeled video data, such as HowToI00M [16], we achieve a remarkable advancement on standard DVC datasets like YouCook2 [31] and ActivityNet [13]. We out-perform the previous state-of-the-art Vid2Seq [27] across a majority of metrics, achieving this with just 0.4% of the unlabeled video data used for pre-training by Vid2Seq.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74676bde-0a86-4627-9f77-16f7e95fc2deCited by top-tier papers14
- Threading Keyframe with Narratives: MLLMs as Strong Long Video ComprehendersBo Fang, Yuxin Song, Haoyuan Sun, Qiangqiang Wu et al.ICLR 2026 · 13 citations
- Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video UnderstandingJialuo Li, Bin Li, Jiahao Li, Yan LuCVPR 2026 · 11 citations
- Exploring Temporal Event Cues for Dense Video Captioning in Cyclic Co-LearningZhuyang Xie, Yan Yang, Yankai Yu, Jie Wang et al.AAAI 2025 · 5 citations
- Enrich and Detect: Video Temporal Grounding With Multimodal LlmsShraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa et al.ICCV 2025 · 4 citations
- HiCM²: Hierarchical Compact Memory Modeling for Dense Video CaptioningMinkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi et al.AAAI 2025 · 2 citations
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng et al.ICCV 2021 · 238 citations
- Watch, Listen and Tell: Multi-Modal Weakly Supervised Dense Event CaptioningTanzila Rahman, Bicheng Xu, Leonid SigalICCV 2019 · 89 citations
- Drop-DTW: Aligning Common Signal Between Sequences While Dropping OutliersNikita Dvornik, Isma Hadji, Konstantinos G. Derpanis, Animesh Garg et al.NeurIPS 2021 · 78 citations
Related papers
- DiffDVC: Accurate Event Detection for Dense Video Captioning via Diffusion ModelsWei Chen, Jianwei Niu, Xuefeng Liu, Zhendong Wang et al.AAAI 2025 · 2 citations
- Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video CaptioningAntoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech et al.CVPR 2023
- Large-Scale Pre-Training for Grounded Video Caption GenerationEvangelos Kazakos, Cordelia Schmid, Josef SivicICCV 2025
- Event-Equalized Dense Video CaptioningKangyi Wu, Pengna Li, Jingwen Fu, Yizhe Li et al.CVPR 2025
- Hierarchical Context-aware Network for Dense Video Event CaptioningLei Ji, Xianglin Guo, Haoyang Huang, Xilin ChenACL 2021
