Visibility Aware Human-Object Interaction Tracking from Single RGB Camera
Xianghui Xie, Bharat Lal Bhatnagar, Gerard Pons-Moll
Abstract
Capturing the interactions between humans and their environment in 3D is important for many applications in robotics, graphics, and vision. Recent works to reconstruct the 3D human and object from a single RGB image do not have consistent relative translation across frames because they assume a fixed depth. Moreover, their performance drops significantly when the object is occluded. In this work, we propose a novel method to track the 3D human, object, contacts, and relative translation across frames from a single RGB camera, while being robust to heavy occlusions. Our method is built on two key insights. First, we condition our neural field reconstructions for human and object on per-frame SMPL model estimates obtained by pre-fitting SMPL to a video sequence. This improves neural reconstruction accuracy and produces coherent relative translation across frames. Second, human and object motion from visible frames provides valuable information to infer the occluded object. We propose a novel transformerbased neural network that explicitly uses object visibility and human motion to leverage neighboring frames to make predictions for the occluded frames. Building on these insights, our method is able to track both human and object robustly even under occlusions. Experiments on two datasets show that our method significantly improves over the state-of-the-art methods. Our code and pretrained models are available at: https://virtualhumans.mpiinf.mpg.de/VisTracker .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0ff1949-9e9b-4ea1-84fd-112ffb3f98fbCited by top-tier papers30
- InterDiff: Generating 3D Human-Object Interactions with Physics-Informed DiffusionSirui Xu, Zhengyuan Li, Yu-Xiong Wang, Liang-Yan GuiICCV 2023 · 201 citations
- Omnigrasp: Grasping Diverse Objects with Simulated HumanoidsZhengyi Luo, Jinkun Cao, Sammy Christen, Alexander Winkler et al.NeurIPS 2024 · 66 citations
- DECO: Dense Estimation of 3D Human-Scene Contact In The WildShashank Tripathi, Agniv Chatterjee, Jean-Claude Passy, Hongwei Yi et al.ICCV 2023 · 54 citations
- Human-3Diffusion: Realistic Avatar Creation via Explicit 3D Consistent Diffusion ModelsYuxuan Xue, Xianghui Xie, Riccardo Marin, Gerard Pons-MollNeurIPS 2024 · 49 citations
- Paint-it: Text-to-Texture Synthesis via Deep Convolutional Texture Map Optimization and Physically-Based RenderingKim Youwang, Tae-Hyun Oh, Gerard Pons-MollCVPR 2024 · 30 citations
Builds on42
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- Neural Unsigned Distance Fields for Implicit Function LearningJulian Chibane, Aymen Mir, Gerard Pons-MollNeurIPS 2020 · 415 citations
Related papers
- RHINO: Reconstructing Human Interactions with Novel Objects from Monocular VideosLixin Xue, Chengwei Zheng, Georgios Paschalidis, Chen Guo et al.CVPR 2026
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri et al.CVPR 2025
- Free-Moving Object Reconstruction and Pose Estimation with Virtual CameraHaixin Shi, Yinlin Hu, Daniel Koguciuk, Juan-Ting Lin et al.AAAI 2025 · 2 citations
- DeepHuman: 3D Human Reconstruction From a Single ImageZerong Zheng, Tao Yu, Yixuan Wei, Qionghai Dai et al.ICCV 2019 · 367 citations
