3D Part Guided Image Editing for Fine-Grained Object Understanding
Zongdai Liu, Feixiang Lu, Peng Wang, Hui Miao, Liangjun Zhang, Ruigang Yang, Bin Zhou
Abstract
Holistically understanding an object with its 3D movable parts is essential for visual models of a robot to interact with the world. For example, only by understanding many possible part dynamics of other vehicles (e.g., door or trunk opening, taillight blinking for changing lane), a self-driving vehicle can be success in dealing with emergency cases. However, existing visual models tackle rarely on these situations, but focus on bounding box detection. In this paper, we fill this important missing piece in autonomous driving by solving two critical issues. First, for dealing with data scarcity, we propose an effective training data generation process by fitting a 3D car model with dynamic parts to cars in real images. This allows us to directly edit the real images using the aligned 3D parts, yielding effective training data for learning robust deep neural networks (DNNs). Secondly, to benchmark the quality of 3D part understanding, we collected a large dataset in real driving scenario with cars in uncommon states (CUS), i.e. with door or trunk opened etc., which demonstrates that our trained network with edited images largely outperforms other baselines in terms of 2D detection and instance segmentation accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8a20843-c2bc-4ff4-8226-0b70727b8324Cited by top-tier papers3
- Understanding Multi-Granularity for Open-Vocabulary Part SegmentationJiho Choi, Seonho Lee, Seungho Lee, Minhyun Lee et al.NeurIPS 2024 · 7 citations
- Robust 2D/3D Vehicle Parsing in Arbitrary Camera Views for CVISHui Miao, Feixiang Lu, Zongdai Liu, Liangjun Zhang et al.ICCV 2021 · 2 citations
- Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part SegmentationJiho Choi, Seonho Lee, Minhyun Lee, Seungho Lee et al.CVPR 2025
Related papers
- Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous DrivingAlexey Nekrasov, Malcolm Burdorf, Stewart Worrall, Bastian Leibe et al.CVPR 2025
- Rethinking Open-World Object Detection in Autonomous Driving ScenariosZeyu Ma, Yang Yang, Guoqing Wang, Xing Xu et al.ACM MM 2022 · 39 citations
- 3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object DetectionAlexander Lehner, Stefano Gasperini, Alvaro Marcos-Ramiro, Michael Schmidt et al.CVPR 2022 · 61 citations
- OpenBox: Annotate Any Bounding Boxes in 3DIn-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio et al.NeurIPS 2025 · 7 citations
- 3DRealCar: An In-the-Wild RGB-D Car Dataset with 360-Degree ViewsXiaobiao Du, Yida Wang, Haiyang Sun, Zhuojie Wu et al.ICCV 2025 · 10 citations
