Learning Action Changes by Measuring Verb-Adverb Textual Relationships
Davide Moltisanti, Frank Keller, Hakan Bilen, Laura Sevilla-Lara
Abstract
The goal of this work is to understand the way actions are performed in videos. That is, given a video, we aim to predict an adverb indicating a modification applied to the action (e.g. cut "finely"). We cast this problem as a regression task. We measure textual relationships between verbs and adverbs to generate a regression target representing the action change we aim to learn. We test our approach on a range of datasets and achieve state-of-the-art results on both adverb prediction and antonym classification. Furthermore, we outperform previous work when we lift two commonly assumed conditions: the availability of action labels during testing and the pairing of adverbs as antonyms. Existing datasets for adverb recognition are either noisy, which makes learning difficult, or contain actions whose appearance is not influenced by adverbs, which makes evaluation less reliable. To address this, we collect a new high quality dataset: Adverbs in Recipes (AIR). We focus on instructional recipes videos, curating a set of actions that exhibit meaningful visual changes when performed differently. Videos in AIR are more tightly trimmed and were manually reviewed by multiple annotators to ensure high labelling quality. Results show that models learn better from AIR given its cleaner videos. At the same time, adverb prediction on AIR is challenging, demonstrating that there is considerable room for improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7be3095-211e-413c-ab38-e245d4f1ae4cCited by top-tier papers3
- Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional VideosKumaranage Ravindu Yasas Nagasinghe, Honglu Zhou, Malitha Gunawardhana, Martin Renqiang Min et al.CVPR 2024 · 5 citations
- Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video UnderstandingKaiting Liu, Hazel DoughtyICLR 2026
- SRTube: Video-Language Pre-Training with Action-Centric Video Tube Features and Semantic Role LabelingJu-Hee Lee, Je-Won KangCVPR 2024
Builds on10
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- TACo: Token-aware Cascade Contrastive Learning for Video-Text AlignmentJianwei Yang, Yonatan Bisk, Jianfeng GaoICCV 2021 · 159 citations
- Further Understanding Videos through Adverbs: A New Video TaskBo Pang, Kaiwen Zha, Yifan Zhang, Cewu LuAAAI 2020 · 18 citations
- How Do You Do It? Fine-Grained Action Understanding with Pseudo-AdverbsHazel Doughty, Cees G. M. SnoekCVPR 2022 · 16 citations
Related papers
- Action Modifiers: Learning From Adverbs in Instructional VideosHazel Doughty, Ivan Laptev, Walterio W. Mayol-Cuevas, Dima DamenCVPR 2020
- Zero-Shot Anticipation for Instructional ActivitiesFadime Sener, Angela YaoICCV 2019 · 75 citations
- GePSAn: Generative Procedure Step Anticipation in Cooking VideosMohamed Ashraf Abdelsalam, Samrudhdhi B. Rangrej, Isma Hadji, Nikita Dvornik et al.ICCV 2023 · 10 citations
- Aligning Actions Across Recipe GraphsLucia Donatelli, Theresa Schmidt, Debanjali Biswas, Arne Köhn et al.EMNLP 2021
- Opening the Vocabulary of Egocentric ActionsDibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela YaoNeurIPS 2023 · 28 citations
