Speech2Action: Cross-Modal Supervision for Action Recognition
Arsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar, Cordelia Schmid, Andrew Zisserman
Abstract
Caption: Hello, it's me Speech2Action classifier [answers] phone Hello, it's me [answers] phone Thanks for calling so soon [answers] phone Hello Dad, are you still there? action: dialogue: action: dialogue: action: dialogue Unlabelled videos She knows he's right. Jane's cell RINGS. She lets it ring again, then answers it. JANE (into phone) Hello, it's me. Movie screenplays Weak label: [answer] phone Figure 1. Weakly Supervised Learning of Actions from Speech Alone: The co-occurrence of speech and scene descriptions in movie screenplays (text) is used to learn a Speech2Action model that predicts actions from transcribed speech alone. Weak labels for visual actions can then be obtained by applying this model to the speech in a large unlabelled set of movies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d01991b6-264d-451c-b012-bea776e549f0Cited by top-tier papers15
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
- Labelling unlabelled videos from scratch with multi-modal self-supervisionYuki Markus Asano, Mandela Patrick, Christian Rupprecht, Andrea VedaldiNeurIPS 2020 · 169 citations
- Verbs in Action: Improving verb understanding in video-language modelsLiliane Momeni, Mathilde Caron, Arsha Nagrani, Andrew Zisserman et al.ICCV 2023 · 93 citations
- CoVR: Learning Composed Video Retrieval from Web Video CaptionsLucas Ventura, Antoine Yang, Cordelia Schmid, Gül VarolAAAI 2024 · 81 citations
- How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the WildOkan Köpüklü, Maja Taseska, Gerhard RigollICCV 2021 · 59 citations
Builds on2
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- Learning to Segment Actions from Observation and NarrationDaniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer et al.ACL 2020 · 24 citations
- Learning Human-Human Interactions in Images from Weak Textual SupervisionMorris Alper, Hadar Averbuch-ElorICCV 2023 · 4 citations
- Learning with Weak Supervision for Email Intent DetectionKai Shu, Subhabrata Mukherjee, Guoqing Zheng, Ahmed Hassan Awadallah et al.SIGIR 2020 · 26 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Set-Constrained Viterbi for Set-Supervised Action SegmentationJun Li, Sinisa TodorovicCVPR 2020
