Action Modifiers: Learning From Adverbs in Instructional Videos
Hazel Doughty, Ivan Laptev, Walterio W. Mayol-Cuevas, Dima Damen
摘要
We present a method to learn a representation for adverbs from instructional videos using weak supervision from the accompanying narrations. Key to our method is the fact that the visual representation of the adverb is highly dependent on the action to which it applies, although the same adverb will modify multiple actions in a similar way. For instance, while spread quickly' and mix quickly' will look dissimilar, we can learn a common representation that allows us to recognize both, among other actions. We formulate this as an embedding problem, and use scaled dot product attention to learn from weakly-supervised video narrations. We jointly learn adverbs as invertible transformations which operate on the embedding space, so as to add or remove the effect of the adverb. As there is no prior work on weakly supervised learning from adverbs, we gather paired action-adverb annotations from a subset of the HowTo100M dataset, for 6 adverbs: quickly/slowly, finely/coarsely and partially/completely. Our method outperforms all baselines for video-to-adverb retrieval with a performance of 0.719 mAP. We also demonstrate our model's ability to attend to the relevant video parts in order to determine the adverb for a given action.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Verbs in Action: Improving verb understanding in video-language modelsLiliane Momeni, Mathilde Caron, Arsha Nagrani, Andrew Zisserman 等ICCV 2023 · 被引用 93 次
- PoseFix: Correcting 3D Human Poses with Natural LanguageGinger Delmas, Philippe Weinzaepfel, Francesc Moreno-Noguer, Grégory RogezICCV 2023 · 被引用 49 次
- Sign Language Video Retrieval with Free-Form Textual QueriesAmanda Cardoso Duarte, Samuel Albanie, Xavier Giró-i-Nieto, Gül VarolCVPR 2022 · 被引用 27 次
- Look for the Change: Learning Object States and State-Modifying Actions from Untrimmed Web VideosTomás Soucek, Jean-Baptiste Alayrac, Antoine Miech, Ivan Laptev 等CVPR 2022 · 被引用 19 次
它引用的顶会 Paper3
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- Fine-Grained Action Retrieval Through Multiple Parts-of-Speech EmbeddingsMichael Wray, Gabriela Csurka, Diane Larlus, Dima DamenICCV 2019 · 被引用 185 次
- Zero-Shot Anticipation for Instructional ActivitiesFadime Sener, Angela YaoICCV 2019 · 被引用 75 次
相关 Paper
- How Do You Do It? Fine-Grained Action Understanding with Pseudo-AdverbsHazel Doughty, Cees G. M. SnoekCVPR 2022 · 被引用 16 次
- Learning Action Changes by Measuring Verb-Adverb Textual RelationshipsDavide Moltisanti, Frank Keller, Hakan Bilen, Laura Sevilla-LaraCVPR 2023
- Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional VideosReuben Tan, Bryan A. Plummer, Kate Saenko, Hailin Jin 等NeurIPS 2021 · 被引用 30 次
- Temporal Alignment Networks for Long-term VideoTengda Han, Weidi Xie, Andrew ZissermanCVPR 2022 · 被引用 60 次
- Opening the Vocabulary of Egocentric ActionsDibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela YaoNeurIPS 2023 · 被引用 28 次
