Lune

CVPR2021Top-tier venue

Look Before You Speak: Visually Contextualized Utterances

Paul Hongsuck Seo, Arsha Nagrani, Cordelia Schmid

2021Year
24Top-tier citations

Abstract

It makes it feel healthier. Now slip that nut back on and screw it down. It's going to take about five minutes. … Transcript: I'm going to go ahead and slip that into place and I'm going to make note of which way the arrow is going in relation to the arrow on our guard. They both need to be going the same direction next. Prediction Next utterance candidates Input Video ✔ Figure 1: Visually Contextualised Future Utterance Prediction. Given an instructional video with paired text and video data, we predict the next utterance in the video using a Co-attentional Multimodal Video Transformer. Our model trained on this task also achieves state-of-the-art performance on downstream VideoQA benchmarks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 0b1656a9-e926-4fb0-b230-ef7d1d9eeb42

Cited by top-tier papers24

Ask how each one uses it

Builds on11

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines