Look for the Change: Learning Object States and State-Modifying Actions from Untrimmed Web Videos
Tomás Soucek, Jean-Baptiste Alayrac, Antoine Miech, Ivan Laptev, Josef Sivic
Abstract
Human actions often induce changes of object states such as "cutting an apple", "cleaning shoes" or "pouring coffee". In this paper, we seek to temporally localize object states (e.g. "empty" and "full" cup) together with the corresponding state-modifying actions ("pouring coffee") in long uncurated videos with minimal supervision. The contributions of this work are threefold. First, we develop a self-supervised model for jointly learning state-modifying actions together with the corresponding object states from an uncurated set of videos from the Internet. The model is self-supervised by the causal ordering signal, i.e. initial object state → manipulating action → end state. Second, to cope with noisy uncurated training data, our model incorporates a noise adaptive weighting module supervised by a small number of annotated still images, that allows to efficiently filter out irrelevant videos during training. Third, we collect a new dataset with more than 2600 hours of video and 34 thousand changes of object states, and manually annotate a part of this data to validate our approach. Our results demonstrate substantial improvements over prior work in both action and object state-recognition in video.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?Qi Zhao, Shijie Wang, Ce Zhang, Changcheng Fu et al.ICLR 2024 · 93 citations
- SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional VideosYulei Niu, Wenliang Guo, Long Chen, Xudong Lin et al.ICLR 2024 · 26 citations
- Chop & Learn: Recognizing and Generating Object-State CompositionsNirat Saini, Hanyu Wang, Archana Swaminathan, Vinoj Jayasundara et al.ICCV 2023 · 20 citations
- Video State-Changing Object SegmentationJiangwei Yu, Xiang Li, Xinran Zhao, Hongming Zhang et al.ICCV 2023 · 16 citations
- Learning Object State Changes in Videos: An Open-World PerspectiveZihui Xue, Kumar Ashutosh, Kristen GraumanCVPR 2024 · 12 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Multiscale Vision TransformersHaoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li et al.ICCV 2021 · 1,611 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- Weakly Supervised Human-Object Interaction Detection in Video via Contrastive Spatiotemporal RegionsShuang Li, Yilun Du, Antonio Torralba, Josef Sivic et al.ICCV 2021 · 17 citations
- Learning State-Aware Visual Representations from Audible InteractionsHimangi Mittal, Pedro Morgado, Unnat Jain, Abhinav GuptaNeurIPS 2022 · 30 citations
- What, When, and Where? Self-Supervised Spatio- Temporal Grounding in Untrimmed Multi-Action Videos from Narrated InstructionsBrian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann et al.CVPR 2024
- Action Shuffle Alternating Learning for Unsupervised Action SegmentationJun Li, Sinisa TodorovicCVPR 2021
- Unsupervised Open-Vocabulary Object Localization in VideosKe Fan, Zechen Bai, Tianjun Xiao, Dominik Zietlow et al.ICCV 2023 · 14 citations
