Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning
Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont-Tuset, Ivan Laptev, Josef Sivic, Cordelia Schmid
Abstract
Dense video captioning Hey guys today I am going to teach you how to ski The kids make it look easy First slope, congratz! Vid2Seq <1s><8s>The man is fastening the dog. <20s><50s>The dogs are pulling the sled. <45s><49s>The man is saying hello.
Figure 1. Vid2Seq is a visual language model that predicts dense event captions together with their temporal grounding in the video by generating a single sequence of tokens (right). This ability is enabled by large-scale pretraining on unlabeled narrated videos (left).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5aeee652-6f5a-4ecf-b21e-6caa4a0186a8Cited by top-tier papers123
- UniVTG: Towards Unified Video-Language Temporal GroundingKevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick et al.ICCV 2023 · 221 citations
- Time-R1: Post-Training Large Vision Language Model for Temporal Video GroundingYe Wang, Ziheng Wang, Boshen Xu, Yang Du et al.NeurIPS 2025 · 143 citations
- Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningLong Qian, Juncheng Li, Yu Wu, Yaobo Ye et al.ICML 2024 · 121 citations
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun et al.CVPR 2024 · 83 citations
- UnLoc: A Unified Framework for Video Localization TasksShen Yan, Xuehan Xiong, Arsha Nagrani, Anurag Arnab et al.ICCV 2023 · 82 citations
Builds on63
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online RefinementHao Wu, Huabin Liu, Yu Qiao, Xiao SunCVPR 2024 · 8 citations
- Streaming Dense Video CaptioningXingyi Zhou, Anurag Arnab, Shyamal Buch, Shen Yan et al.CVPR 2024 · 33 citations
- Sketch, Ground, and Refine: Top-Down Dense Video CaptioningChaorui Deng, Shizhe Chen, Da Chen, Yuan He et al.CVPR 2021
- VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video CaptioningJi Soo Lee, Jongha Kim, Jeehye Na, Jinyoung Park et al.AAAI 2025 · 11 citations
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng et al.ICCV 2021 · 238 citations
