Occlusion-Aware Temporally Consistent Amodal Completion for 3D Human-Object Interaction Reconstruction
Hyungjun Doh, Dong In Lee, Seunggeun Chi, Pin-Hao Huang, Kwonjoon Lee, Sangpil Kim, Karthik Ramani
Abstract
We introduce a novel framework for reconstructing dynamic human-object interactions from monocular video that overcomes challenges associated with occlusions and temporal inconsistencies. Traditional 3D reconstruction methods typically assume static objects or full visibility of dynamic subjects, leading to degraded performance when these assumptions are violated-particularly in scenarios where mutual occlusions occur. To address this, our framework leverages amodal completion to infer the complete structure of partially obscured regions. Unlike conventional approaches that operate on individual frames, our method integrates temporal context, enforcing coherence across video sequences to incrementally refine and stabilize reconstructions. This template-free strategy adapts to varying conditions without relying on predefined models, significantly enhancing the recovery of intricate details in dynamic scenes. We validate our approach using 3D Gaussian Splatting on challenging monocular videos, demonstrating superior precision in handling occlusions and maintaining temporal stability compared to existing techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af1c0f94-7a65-4a21-8bda-e4efc7233e24Cited by top-tier papers1
Ask how each one uses itBuilds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields ReconstructionCheng Sun, Min Sun, Hwann-Tzong ChenCVPR 2022 · 859 citations
Related papers
- Illumination-Consistent Human-Scene Reconstruction from Monocular VideoRongbin Zheng, Wensheng Li, Lingzhe Zeng, Dong Wang et al.CVPR 2026
- HAIF-GS: Hierarchical and Induced Flow-Guided Gaussian Splatting for Dynamic SceneJianing Chen, Zehao Li, Yujun Cai, Hao Jiang et al.NeurIPS 2025 · 14 citations
- MOSAIC-GS: Monocular Scene Reconstruction via Advanced Initialization for Complex Dynamic EnvironmentsSvitlana Morkva, Vaishakh Patil, Alessio Tonioni, Michael Oechsle et al.CVPR 2026
- Secondary Motion-Aware 3D Clothed Gaussian Avatars from Monocular VideosSeungeun Lee, Seungjun Moon, Hah Min Lew, Ji-Su Kang et al.ICLR 2026
- Uncertainty Matters in Dynamic Gaussian Splatting for Monocular 4D ReconstructionFengzhi Guo, Chih-Chuan Hsu, Sihao Ding, Cheng ZhangICLR 2026 · 6 citations
