WALT3D: Generating Realistic Training Data from Time-Lapse Imagery for Reconstructing Dynamic Objects Under Occlusion
Khiem Vuong, N. Dinesh Reddy, Robert Tamburo, Srinivasa G. Narasimhan
Abstract
Current methods for 2D and 3D object understanding struggle with severe occlusions in busy urban environments, partly due to the lack of large-scale labeled ground-truth annotations for learning occlusion. In this work, we introduce a novel framework for automatically generating a large, realistic dataset of dynamic objects under occlusions using freely available time-lapse imagery. By leveraging off-the-shelf2D (bounding box, segmentation, keypoint) and 3D (pose, shape) predictions as pseudo-groundtruth, unoccluded 3D objects are identified automatically and composited into the background in a clip-art style, ensuring realistic appearances and physically accurate occlusion configurations. The resulting clip-art image with pseudogroundtruth enables efficient training of object reconstruction methods that are robust to occlusions. Our method demonstrates significant improvements in both 2D and 3D reconstruction, particularly in scenarios with heavily occluded objects like vehicles and people in urban scenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b66356b7-b595-4dcc-b224-be2b7ba725f8Cited by top-tier papers3
- ROADWork: A Dataset and Benchmark for Learning to Recognize, Observe, Analyze and Drive Through Work ZonesAnurag Ghosh, Shen Zheng, Robert Tamburo, Khiem Vuong et al.ICCV 2025 · 4 citations
- Stable Diffusion-Based Approach for Human De-OcclusionSeung Young Noh, Ju Yong ChangACM MM 2025
- PHAC: Promptable Human Amodal CompletionSeung Young Noh, Ju Yong ChangCVPR 2026
Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- Humans in 4D: Reconstructing and Tracking Humans with TransformersShubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa et al.ICCV 2023 · 390 citations
- Putting People in their Place: Monocular Regression of 3D People in DepthYu Sun, Wu Liu, Qian Bao, Yili Fu et al.CVPR 2022 · 152 citations
Related papers
- WALT: Watch And Learn 2D amodal representation from Time-lapse imageryN. Dinesh Reddy, Robert Tamburo, Srinivasa G. NarasimhanCVPR 2022 · 25 citations
- ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic GraspingShun Iwase, Muhammad Zubair Irshad, Katherine Liu, Vitor Guizilini et al.CVPR 2025
- VOccl3D: A Video Benchmark Dataset for 3D Human Pose and Shape Estimation Under Real OcclusionsYash Garg, Saketh Bachu, Arindam Dutta, Rohit Lal et al.ICCV 2025
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic CamerasYe Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani et al.CVPR 2022 · 111 citations
- Deep 3D Mask Volume for View Synthesis of Dynamic ScenesKai-En Lin, Lei Xiao, Feng Liu, Guowei Yang et al.ICCV 2021 · 42 citations
