Speech2Action: Cross-Modal Supervision for Action Recognition
Arsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar, Cordelia Schmid, Andrew Zisserman
摘要
Caption: Hello, it's me Speech2Action classifier [answers] phone Hello, it's me [answers] phone Thanks for calling so soon [answers] phone Hello Dad, are you still there? action: dialogue: action: dialogue: action: dialogue Unlabelled videos She knows he's right. Jane's cell RINGS. She lets it ring again, then answers it. JANE (into phone) Hello, it's me. Movie screenplays Weak label: [answer] phone Figure 1. Weakly Supervised Learning of Actions from Speech Alone: The co-occurrence of speech and scene descriptions in movie screenplays (text) is used to learn a Speech2Action model that predicts actions from transcribed speech alone. Weak labels for visual actions can then be obtained by applying this model to the speech in a large unlabelled set of movies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang 等CVPR 2022 · 被引用 264 次
- Labelling unlabelled videos from scratch with multi-modal self-supervisionYuki Markus Asano, Mandela Patrick, Christian Rupprecht, Andrea VedaldiNeurIPS 2020 · 被引用 169 次
- Verbs in Action: Improving verb understanding in video-language modelsLiliane Momeni, Mathilde Caron, Arsha Nagrani, Andrew Zisserman 等ICCV 2023 · 被引用 93 次
- CoVR: Learning Composed Video Retrieval from Web Video CaptionsLucas Ventura, Antoine Yang, Cordelia Schmid, Gül VarolAAAI 2024 · 被引用 81 次
- How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the WildOkan Köpüklü, Maja Taseska, Gerhard RigollICCV 2021 · 被引用 59 次
它引用的顶会 Paper2
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
相关 Paper
- Learning to Segment Actions from Observation and NarrationDaniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer 等ACL 2020 · 被引用 24 次
- Learning Human-Human Interactions in Images from Weak Textual SupervisionMorris Alper, Hadar Averbuch-ElorICCV 2023 · 被引用 4 次
- Learning with Weak Supervision for Email Intent DetectionKai Shu, Subhabrata Mukherjee, Guoqing Zheng, Ahmed Hassan Awadallah 等SIGIR 2020 · 被引用 26 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Set-Constrained Viterbi for Set-Supervised Action SegmentationJun Li, Sinisa TodorovicCVPR 2020
