Weakly Supervised Human-Object Interaction Detection in Video via Contrastive Spatiotemporal Regions
Shuang Li, Yilun Du, Antonio Torralba, Josef Sivic, Bryan C. Russell
Abstract
We introduce the task of weakly supervised learning for detecting human and object interactions in videos. Our task poses unique challenges as a system does not know what types of human-object interactions are present in a video or the actual spatiotemporal location of the human and the object. To address these challenges, we introduce a contrastive weakly supervised training loss that aims to jointly associate spatiotemporal regions in a video with an action and object vocabulary and encourage temporal continuity of the visual appearance of moving objects as a form of self-supervision. To train our model, we introduce a dataset comprising over 6.5k videos with human-object interaction annotations that have been semi-automatically curated from sentence captions associated with the videos. We demonstrate improved performance over weakly supervised baselines adapted to our task on our video dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c25573ac-545a-4463-a2b4-7c453a89c8f0Cited by top-tier papers7
- Learning Transferable Human-Object Interaction Detector with Natural Language SupervisionSuchen Wang, Yueqi Duan, Henghui Ding, Yap-Peng Tan et al.CVPR 2022 · 66 citations
- Helping Hands: An Object-Aware Ego-Centric Video Recognition ModelChuhan Zhang, Ankush Gupta, Andrew ZissermanICCV 2023 · 39 citations
- Open Set Video HOI detection from Action-centric Chain-of-Look PromptingNan Xi, Jingjing Meng, Junsong YuanICCV 2023 · 10 citations
- MATRIX: Mask Track Alignment for Interaction-aware Video GenerationSiyoon Jin, Seongchan Kim, Jae Ho Lee, Dahyun Chung et al.ICLR 2026 · 4 citations
- What, When, and Where? Self-Supervised Spatio- Temporal Grounding in Untrimmed Multi-Action Videos from Narrated InstructionsBrian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann et al.CVPR 2024
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- SCT: Set Constrained Temporal Transformer for Set Supervised Action SegmentationMohsen Fayyaz, Jürgen GallCVPR 2020
- Look for the Change: Learning Object States and State-Modifying Actions from Untrimmed Web VideosTomás Soucek, Jean-Baptiste Alayrac, Antoine Miech, Ivan Laptev et al.CVPR 2022 · 19 citations
- Exploring Heterogeneous Clues for Weakly-Supervised Audio-Visual Video ParsingYu Wu, Yi YangCVPR 2021
- Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action LocalizationGeuntaek Lim, Hyunwoo Kim, Joonsoo Kim, Yukyung ChoiACM MM 2024 · 11 citations
- Opening the Vocabulary of Egocentric ActionsDibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela YaoNeurIPS 2023 · 28 citations
