Lune

CVPR2020顶会

Speech2Action: Cross-Modal Supervision for Action Recognition

Arsha Nagrani, Chen Sun, David Ross, Rahul Sukthankar, Cordelia Schmid, Andrew Zisserman

2020年份
15顶会引用

摘要

Caption: Hello, it's me Speech2Action classifier [answers] phone Hello, it's me [answers] phone Thanks for calling so soon [answers] phone Hello Dad, are you still there? action: dialogue: action: dialogue: action: dialogue Unlabelled videos She knows he's right. Jane's cell RINGS. She lets it ring again, then answers it. JANE (into phone) Hello, it's me. Movie screenplays Weak label: [answer] phone Figure 1. Weakly Supervised Learning of Actions from Speech Alone: The co-occurrence of speech and scene descriptions in movie screenplays (text) is used to learn a Speech2Action model that predicts actions from transcribed speech alone. Weak labels for visual actions can then be obtained by applying this model to the speech in a large unlabelled set of movies.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper15

问问它们各自怎么用它

它引用的顶会 Paper2

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖