Capturing Closely Interacted Two-Person Motions with Reaction Priors
Qi Fang, Yinghui Fan, Yanjun Li, Junting Dong, Dingwei Wu, Weidong Zhang, Kang Chen
摘要
In this paper, we focus on capturing closely interacted two-person motions from monocular videos, an important yet understudied topic. Unlike less-interacted motions, closely interacted motions contain frequently occurring inter-human occlusions, which pose significant challenges to existing capturing algorithms. To address this problem, our key observation is that close physical interactions between two subjects typically happen under very specific situations (e.g., handshake, hug, etc.), and such situational contexts contain strong prior semantics to help infer the poses of occluded joints. In this spirit, we introduce reaction priors, which are invertible neural networks that bi-directionally model the pose probability distributions of one person given the pose of the other. The learned reaction priors are then incorporated into a querybased pose estimator, which is a decoder-only Transformer with self-attentions on both intra-joint and inter-joint relationships. We demonstrate that our design achieves considerably higher performance than previous methods on multiple benchmarks. What's more, as existing datasets lack sufficient cases of close human-human interactions, we also build a new dataset called Dual-Human to better evaluate different methods. Dual-Human contains around 2k sequences of closely interacted two-person motions, each with synthetic multi-view renderings, contact annotations, and text descriptions. We believe that this new public dataset can significantly promote further research in this area. Our project page is at https://neteasegameai.github.io/Dual-Human/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Ponimator: Unfolding Interactive Pose for Versatile Human-Human Interaction AnimationShaowei Liu, Chuan Guo, Bing Zhou, Jian WangICCV 2025 · 被引用 2 次
- Generating Attribute-Aware Human Motions from Textual PromptXinghan Wang, Kun Xu, Fei Li, Cao Sheng 等AAAI 2026
- MAMMA: Markerless Accurate Multi-person Motion AcquisitionHanz Cuevas Velasquez, Anastasios Yiannakidis, Soyong Shin, Giorgio Becherini 等CVPR 2026
- Reconstructing Close Human Interaction with Appearance and Proxemics ReasoningBuzhen Huang, Chen Li, Chongyang Xu, Dongyue Lu 等CVPR 2025
它引用的顶会 Paper40
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 被引用 672 次
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 被引用 509 次
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang 等CVPR 2022 · 被引用 462 次
- HuMoR: 3D Human Motion Model for Robust Pose EstimationDavis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang 等ICCV 2021 · 被引用 398 次
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa 等ICCV 2023 · 被引用 390 次
相关 Paper
- Closely Interactive Human Reconstruction with Proxemics and Physics-Guided AdaptionBuzhen Huang, Chen Li, Chongyang Xu, Liang Pan 等CVPR 2024
- A Dual-Masked Auto-Encoder for Robust Motion Capture with Spatial-Temporal Skeletal Token CompletionJunkun Jiang, Jie Chen, Yike GuoACM MM 2022 · 被引用 8 次
- Multi-Person Extreme Motion PredictionWen Guo, Xiaoyu Bie, Xavier Alameda-Pineda, Francesc Moreno-NoguerCVPR 2022 · 被引用 64 次
- CrossHOI: Learning Cross-View Representations for Monocular 3D Human-Object Interaction ReconstructionPei Geng, Shanshan Zhang, Jian YangCVPR 2026
- End-to-End Detection and Pose Estimation of Two Interacting HandsDonguk Kim, Kwang In Kim, Seungryul BaekICCV 2021 · 被引用 57 次
