Interacted Object Grounding in Spatio-Temporal Human-Object Interactions
Xiaoyang Liu, Boran Wen, Xinpeng Liu, Zizheng Zhou, Hongwei Fan, Cewu Lu, Lizhuang Ma, Yulong Chen, Yong-Lu Li
Abstract
Spatio-temporal Human-Object Interaction (ST-HOI) understanding aims at detecting HOIs from videos, which is crucial for activity understanding. However, existing whole-bodyobject interaction video benchmarks overlook the truth that open-world objects are diverse, that is, they usually provide limited and predefined object classes. Therefore, we introduce a new open-world benchmark: Grounding Interacted Objects (GIO) including 1,098 interacted objects class and 290K interacted object boxes annotation. Accordingly, an object grounding task is proposed expecting vision systems to discover interacted objects. Even though today's detectors and grounding methods have succeeded greatly, they perform unsatisfactorily in localizing diverse and rare objects in GIO. This profoundly reveals the limitations of current vision systems and poses a great challenge. Thus, we explore leveraging spatio-temporal cues to address object grounding and propose a 4D question-answering framework (4D-QA) to discover interacted objects from diverse videos. Our method demonstrates significant superiority in extensive experiments compared to current baselines. Data and code will be publicly available at https://github.com/DirtyHarryLYL/HAKE-AVA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47bb19b7-d8a1-45be-bcad-1536d8a6ffb9Cited by top-tier papers5
- HAODiff: Human-Aware One-Step Diffusion via Dual-Prompt GuidanceJue Gong, Tingyu Yang, Jingkai Wang, Zheng Chen et al.NeurIPS 2025 · 5 citations
- FideDiff: Efficient Diffusion Model for High-Fidelity Image Motion DeblurringXiaoyang Liu, Zhengyan Zhou, Zihang Xu, Jiezhang Cao et al.ICLR 2026 · 5 citations
- Reconstructing In-the-Wild Open-Vocabulary Human-Object InteractionsBoran Wen, Dingbang Huang, Zichen Zhang, Jiahong Zhou et al.CVPR 2025
- Human Body Restoration with One-Step Diffusion Model and A New BenchmarkJue Gong, Jingkai Wang, Zheng Chen, Xin Liu et al.ICML 2025
- Homogeneous Dynamics Space for Heterogeneous HumansXinpeng Liu, Junxuan Liang, Chenshuo Zhang, Zixuan Cai et al.CVPR 2025
Builds on21
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
- DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world DetectionLewei Yao, Jianhua Han, Youpeng Wen, Xiaodan Liang et al.NeurIPS 2022 · 285 citations
- Efficient Learning on Point Clouds With Basis Point SetsSergey Prokudin, Christoph Lassner, Javier RomeroICCV 2019 · 156 citations
- HOI Analysis: Integrating and Decomposing Human-Object InteractionYong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li et al.NeurIPS 2020 · 152 citations
Related papers
- Toward Open-Set Human Object Interaction DetectionMingrui Wu, Yuqi Liu, Jiayi Ji, Xiaoshuai Sun et al.AAAI 2024 · 12 citations
- Towards Open-vocabulary HOI Detection with Calibrated Vision-language Models and Locality-aware QueriesZhenhao Yang, Xin Liu, Deqiang Ouyang, Guiduo Duan et al.ACM MM 2024 · 5 citations
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction DynamicsMasatoshi Tateno, Gido Kato, Hirokatsu Kataoka, Yoichi Sato et al.CVPR 2026 · 2 citations
- OmniSTVG: Toward Spatio-Temporal Omni-Object Video GroundingJiali Yao, Xin Gu, Xinran Deng, Mengrui Dai et al.ICLR 2026 · 8 citations
- Human-Object-Object Interaction: Towards Human-Centric Complex Interaction DetectionMingxuan Zhang, Xiao Wu, Zhaoquan Yuan, Qi He et al.ACM MM 2023 · 6 citations
