3D Copy-Paste: Physically Plausible Object Insertion for Monocular 3D Detection
Yunhao Ge, Hong-Xing Yu, Cheng Zhao, Yuliang Guo, Xinyu Huang, Liu Ren, Laurent Itti, Jiajun Wu
Abstract
A major challenge in monocular 3D object detection is the limited diversity and quantity of objects in real datasets. While augmenting real scenes with virtual objects holds promise to improve both the diversity and quantity of the objects, it remains elusive due to the lack of an effective 3D object insertion method in complex real captured scenes. In this work, we study augmenting complex real indoor scenes with virtual objects for monocular 3D object detection. The main challenge is to automatically identify plausible physical properties for virtual assets (e.g., locations, appearances, sizes, etc.) in cluttered real scenes. To address this challenge, we propose a physically plausible indoor 3D object insertion approach to automatically copy virtual objects and paste them into real scenes. The resulting objects in scenes have 3D bounding boxes with plausible physical locations and appearances. In particular, our method first identifies physically feasible locations and poses for the inserted objects to prevent collisions with the existing room layout. Subsequently, it estimates spatially-varying illumination for the insertion location, enabling the immersive blending of the virtual objects into the original scene with plausible appearances and cast shadows. We show that our augmentation method significantly improves existing monocular 3D object models and achieves state-of-the-art performance. For the first time, we demonstrate that a physically plausible 3D object insertion, serving as a generative data augmentation technique, can lead to significant improvements for discriminative downstream tasks such as monocular 3D object detection. Project website: https://gyhandy.github. io/3D-Copy-Paste/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e18d398-b969-4e8d-aab4-8fa0edfda8eeCited by top-tier papers7
- AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based ReferringXinyi Wang, Na Zhao, Zhiyuan Han, Dan Guo et al.AAAI 2025 · 12 citations
- LabelAny3D: Label Any Object 3D in the WildJin Yao, Radowan Mahmud Redoy, Sebastian G. Elbaum, Matthew Dwyer et al.NeurIPS 2025 · 8 citations
- Navigating Beyond Instructions: Vision-and-Language Navigation in Obstructed EnvironmentsHaodong Hong, Sen Wang, Zi Huang, Qi Wu et al.ACM MM 2024 · 4 citations
- Zero-Shot Depth Aware Image Editing With Diffusion ModelsRishubh Parihar, Sachidanand VS, R. Venkatesh BabuICCV 2025 · 3 citations
- MonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular DetectionRishubh Parihar, Srinjay Sarkar, Sarthak Vora, Jogendra Nath Kundu et al.CVPR 2025
Builds on12
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 542 citations
- Disentangling Monocular 3D Object DetectionAndrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Manuel Lopez-Antequera et al.ICCV 2019 · 504 citations
- Deep Parametric Indoor Lighting EstimationMarc-André Gardner, Yannick Hold-Geoffroy, Kalyan Sunkavalli, Christian Gagné et al.ICCV 2019 · 155 citations
- 3D Common Corruptions and Data AugmentationOguzhan Fatih Kar, Teresa Yeo, Andrei Atanov, Amir ZamirCVPR 2022 · 80 citations
- Exploring Geometric Consistency for Monocular 3D Object DetectionQing Lian, Botao Ye, Ruijia Xu, Weilong Yao et al.CVPR 2022 · 34 citations
Related papers
- Partially Fake it Till you Make It: Mixing Real and Fake Thermal Images for Improved Object DetectionFrancesco Bongini, Lorenzo Berlincioni, Marco Bertini, Alberto Del BimboACM MM 2021 · 17 citations
- LiDAR-Aug: A General Rendering-Based Augmentation Framework for 3D Object DetectionJin Fang, Xinxin Zuo, Dingfu Zhou, Shengze Jin et al.CVPR 2021
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object DetectionYung-Hsu Yang, Luigi Piccinelli, Mattia Segù, Siyuan Li et al.ICCV 2025 · 2 citations
- Back to Reality: Weakly-supervised 3D Object Detection with Shape-guided Label EnhancementXiuwei Xu, Yifan Wang, Yu Zheng, Yongming Rao et al.CVPR 2022 · 25 citations
- MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated LabelJunyoung Jung, Seokwon Kim, Jung Uk KimCVPR 2026 · 1 citation
