Interacted Object Grounding in Spatio-Temporal Human-Object Interactions
Xiaoyang Liu, Boran Wen, Xinpeng Liu, Zizheng Zhou, Hongwei Fan, Cewu Lu, Lizhuang Ma, Yulong Chen, Yong-Lu Li
摘要
Spatio-temporal Human-Object Interaction (ST-HOI) understanding aims at detecting HOIs from videos, which is crucial for activity understanding. However, existing whole-bodyobject interaction video benchmarks overlook the truth that open-world objects are diverse, that is, they usually provide limited and predefined object classes. Therefore, we introduce a new open-world benchmark: Grounding Interacted Objects (GIO) including 1,098 interacted objects class and 290K interacted object boxes annotation. Accordingly, an object grounding task is proposed expecting vision systems to discover interacted objects. Even though today's detectors and grounding methods have succeeded greatly, they perform unsatisfactorily in localizing diverse and rare objects in GIO. This profoundly reveals the limitations of current vision systems and poses a great challenge. Thus, we explore leveraging spatio-temporal cues to address object grounding and propose a 4D question-answering framework (4D-QA) to discover interacted objects from diverse videos. Our method demonstrates significant superiority in extensive experiments compared to current baselines. Data and code will be publicly available at https://github.com/DirtyHarryLYL/HAKE-AVA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- HAODiff: Human-Aware One-Step Diffusion via Dual-Prompt GuidanceJue Gong, Tingyu Yang, Jingkai Wang, Zheng Chen 等NeurIPS 2025 · 被引用 5 次
- FideDiff: Efficient Diffusion Model for High-Fidelity Image Motion DeblurringXiaoyang Liu, Zhengyan Zhou, Zihang Xu, Jiezhang Cao 等ICLR 2026 · 被引用 5 次
- Reconstructing In-the-Wild Open-Vocabulary Human-Object InteractionsBoran Wen, Dingbang Huang, Zichen Zhang, Jiahong Zhou 等CVPR 2025
- Human Body Restoration with One-Step Diffusion Model and A New BenchmarkJue Gong, Jingkai Wang, Zheng Chen, Xin Liu 等ICML 2025
- Homogeneous Dynamics Space for Heterogeneous HumansXinpeng Liu, Junxuan Liang, Chenshuo Zhang, Zixuan Cai 等CVPR 2025
它引用的顶会 Paper21
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang 等ICLR 2023 · 被引用 753 次
- DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world DetectionLewei Yao, Jianhua Han, Youpeng Wen, Xiaodan Liang 等NeurIPS 2022 · 被引用 285 次
- Efficient Learning on Point Clouds With Basis Point SetsSergey Prokudin, Christoph Lassner, Javier RomeroICCV 2019 · 被引用 156 次
- HOI Analysis: Integrating and Decomposing Human-Object InteractionYong-Lu Li, Xinpeng Liu, Xiaoqian Wu, Yizhuo Li 等NeurIPS 2020 · 被引用 152 次
相关 Paper
- Toward Open-Set Human Object Interaction DetectionMingrui Wu, Yuqi Liu, Jiayi Ji, Xiaoshuai Sun 等AAAI 2024 · 被引用 12 次
- Towards Open-vocabulary HOI Detection with Calibrated Vision-language Models and Locality-aware QueriesZhenhao Yang, Xin Liu, Deqiang Ouyang, Guiduo Duan 等ACM MM 2024 · 被引用 5 次
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction DynamicsMasatoshi Tateno, Gido Kato, Hirokatsu Kataoka, Yoichi Sato 等CVPR 2026 · 被引用 2 次
- OmniSTVG: Toward Spatio-Temporal Omni-Object Video GroundingJiali Yao, Xin Gu, Xinran Deng, Mengrui Dai 等ICLR 2026 · 被引用 8 次
- Human-Object-Object Interaction: Towards Human-Centric Complex Interaction DetectionMingxuan Zhang, Xiao Wu, Zhaoquan Yuan, Qi He 等ACM MM 2023 · 被引用 6 次
