Lune

CVPR2023Top-tier venue

Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning

Antoine Yang, Arsha Nagrani, Paul Hongsuck Seo, Antoine Miech, Jordi Pont-Tuset, Ivan Laptev, Josef Sivic, Cordelia Schmid

2023Year
123Top-tier citations

Abstract

Dense video captioning Hey guys today I am going to teach you how to ski The kids make it look easy First slope, congratz! Vid2Seq <1s><8s>The man is fastening the dog. <20s><50s>The dogs are pulling the sled. <45s><49s>The man is saying hello.

Figure 1. Vid2Seq is a visual language model that predicts dense event captions together with their temporal grounding in the video by generating a single sequence of tokens (right). This ability is enabled by large-scale pretraining on unlabeled narrated videos (left).

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 5aeee652-6f5a-4ecf-b21e-6caa4a0186a8

Cited by top-tier papers123

Ask how each one uses it

Builds on63

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines