How Do You Do It? Fine-Grained Action Understanding with Pseudo-Adverbs
Hazel Doughty, Cees G. M. Snoek
摘要
We aim to understand how actions are performed and identify subtle differences, such as ‘fold firmly’ vs. ‘fold gently’. To this end, we propose a method which recognizes adverbs across different actions. However, such fine-grained annotations are difficult to obtain and their long-tailed nature makes it challenging to recognize adverbs in rare action-adverb compositions. Our approach therefore uses semi-supervised learning with multiple adverb pseudo-labels to leverage videos with only action labels. Combined with adaptive thresholding of these pseudo-adverbs we are able to make efficient use of the available data while tackling the long-tailed distribution. Additionally, we gather adverb annotations for three existing video retrieval datasets, which allows us to introduce the new tasks of recognizing adverbs in unseen action-adverb compositions and unseen domains. Experiments demonstrate the effectiveness of our method, which outperforms prior work in recognizing adverbs and semi-supervised works adapted for adverb recognition. We also show how adverbs can relate fine-grained actions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Verbs in Action: Improving verb understanding in video-language modelsLiliane Momeni, Mathilde Caron, Arsha Nagrani, Andrew Zisserman 等ICCV 2023 · 被引用 93 次
- FineParser: A Fine-Grained Spatio-Temporal Action Parser for Human-Centric Action Quality AssessmentJinglin Xu, Sibo Yin, Guohao Zhao, Zishuo Wang 等CVPR 2024 · 被引用 31 次
- Learning State-Aware Visual Representations from Audible InteractionsHimangi Mittal, Pedro Morgado, Unnat Jain, Abhinav GuptaNeurIPS 2022 · 被引用 30 次
- Tubelet-Contrastive Self-Supervision for Video-Efficient GeneralizationFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekICCV 2023 · 被引用 13 次
- Training-Free Personalization via Retrieval and Reasoning on FingerprintsDeepayan Das, Davide Talon, Yiming Wang, Massimiliano Mancini 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper29
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
相关 Paper
- Action Modifiers: Learning From Adverbs in Instructional VideosHazel Doughty, Ivan Laptev, Walterio W. Mayol-Cuevas, Dima DamenCVPR 2020
- Learning Action Changes by Measuring Verb-Adverb Textual RelationshipsDavide Moltisanti, Frank Keller, Hakan Bilen, Laura Sevilla-LaraCVPR 2023
- Semi-supervised Learning for Multi-label Video Action DetectionHongcheng Zhang, Xu Zhao, Dongqi WangACM MM 2022 · 被引用 10 次
- Further Understanding Videos through Adverbs: A New Video TaskBo Pang, Kaiwen Zha, Yifan Zhang, Cewu LuAAAI 2020 · 被引用 18 次
- Storyboard-guided Alignment for Fine-grained Video Action RecognitionEnqi Liu, Liyuan Pan, Yan Yang, Yiran Zhong 等NeurIPS 2025 · 被引用 3 次
