Action Sequence Augmentation for Action Anticipation
Yihui Qiu, Deepu Rajan
Abstract
Action anticipation models require an understanding of temporal action patterns and dependencies to predict future actions from previous events. The key challenges arise from the vast number of possible action sequences, given the flexibility in action ordering and the interleaving of multiple goals. Since only a subset of such action sequences are present in action anticipation datasets, there is an inherent ordering bias in them. Another challenge is the presence of noisy input to the models due to erroneous action recognition or other upstream tasks. This paper addresses these challenges by introducing a novel data augmentation strategy that separately augments observed action sequences and next actions. To address biased action ordering, we introduce a grammar induction algorithm that derives a powerful context-free grammar from action sequence data. We also develop an efficient parser to generate plausible next-action candidates beyond the ground truth. For noisy input, we enhance model robustness by randomly deleting or replacing actions in observed sequences. Our experiments on the 50Salads, EGTEA Gaze+, and Epic-Kitchens-100 datasets demonstrate significant performance improvements over existing state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 560f8b2b-e8a1-4baa-805d-ab15cedebe0aCited by top-tier papers1
Ask how each one uses itBuilds on9
- Keeping Your Eye on the Ball: Trajectory Attention in Video TransformersMandela Patrick, Dylan Campbell, Yuki M. Asano, Ishan Misra et al.NeurIPS 2021 · 382 citations
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 270 citations
- Activity Grammars for Temporal Action SegmentationDayoung Gong, Joonseok Lee, Deunsol Jung, Suha Kwak et al.NeurIPS 2023 · 17 citations
- Can't make an Omelette without Breaking some Eggs: Plausible Action Anticipation using Large Video-Language ModelsHimangi Mittal, Nakul Agarwal, Shao-Yuan Lo, Kwonjoon LeeCVPR 2024 · 14 citations
- Context-Aware Sequence Alignment using 4D Skeletal AugmentationTaein Kwon, Bugra Tekin, Siyu Tang, Marc PollefeysCVPR 2022 · 6 citations
Related papers
- AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?Qi Zhao, Shijie Wang, Ce Zhang, Changcheng Fu et al.ICLR 2024 · 93 citations
- Uncertainty-aware Action Decoupling Transformer for Action AnticipationHongji Guo, Nakul Agarwal, Shao-Yuan Lo, Kwonjoon Lee et al.CVPR 2024
- Zero-Shot Anticipation for Instructional ActivitiesFadime Sener, Angela YaoICCV 2019 · 75 citations
- Opening the Vocabulary of Egocentric ActionsDibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela YaoNeurIPS 2023 · 28 citations
- Markov Game Video Augmentation for Action SegmentationNicolas Aziere, Sinisa TodorovicICCV 2023 · 5 citations
