Lune

CVPR2020Top-tier venue

Speech2Action: Cross-Modal Supervision for Action Recognition

Arsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar, Cordelia Schmid, Andrew Zisserman

2020Year
15Top-tier citations

Abstract

Caption: Hello, it's me Speech2Action classifier [answers] phone Hello, it's me [answers] phone Thanks for calling so soon [answers] phone Hello Dad, are you still there? action: dialogue: action: dialogue: action: dialogue Unlabelled videos She knows he's right. Jane's cell RINGS. She lets it ring again, then answers it. JANE (into phone) Hello, it's me. Movie screenplays Weak label: [answer] phone Figure 1. Weakly Supervised Learning of Actions from Speech Alone: The co-occurrence of speech and scene descriptions in movie screenplays (text) is used to learn a Speech2Action model that predicts actions from transcribed speech alone. Weak labels for visual actions can then be obtained by applying this model to the speech in a large unlabelled set of movies.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d01991b6-264d-451c-b012-bea776e549f0

Cited by top-tier papers15

Ask how each one uses it

Builds on2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines