Lune

CVPR2024Top-tier venue

DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement

Hao Wu, Huabin Liu, Yu Qiao, Xiao Sun

2024Year
8Citations
14Top-tier citations

Abstract

We present Dive Into the BoundarieS (DIBS), a novel pretraining framework for dense video captioning (DVC), that elaborates on improving the quality of the generated event captions and their associated pseudo event bound-aries from unlabeled videos. By leveraging the capabil-ities of diverse large language models (LLMs), we gen-erate rich DVC-oriented caption candidates and optimize the corresponding pseudo boundaries under several metic-ulously designed objectives, considering diversity, event-centricity, temporal ordering, and coherence. Moreover, we further introduce a novel online boundary refinement strat-egy that iteratively improves the quality of pseudo bound-aries during training. Comprehensive experiments have been conducted to examine the effectiveness of the pro-posed technique components. By leveraging a substantial amount of unlabeled video data, such as HowToI00M [16], we achieve a remarkable advancement on standard DVC datasets like YouCook2 [31] and ActivityNet [13]. We out-perform the previous state-of-the-art Vid2Seq [27] across a majority of metrics, achieving this with just 0.4% of the unlabeled video data used for pre-training by Vid2Seq.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 74676bde-0a86-4627-9f77-16f7e95fc2de

Cited by top-tier papers14

Ask how each one uses it

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines