LabelAny3D: Label Any Object 3D in the Wild
Jin Yao, Radowan Mahmud Redoy, Sebastian G. Elbaum, Matthew Dwyer, Zezhou Cheng
摘要
Detecting objects in 3D space from monocular input is crucial for applications ranging from robotics to scene understanding. Despite advanced performance in the indoor and autonomous driving domains, existing monocular 3D detection models struggle with in-the-wild images due to the lack of 3D in-the-wild datasets and the challenges of 3D annotation. We introduce LabelAny3D, an analysis-by-synthesis framework that reconstructs holistic 3D scenes from 2D images to efficiently produce high-quality 3D bounding box annotations. Built on this pipeline, we present COCO3D, a new benchmark for open-vocabulary monocular 3D detection, derived from the MS-COCO dataset and covering a wide range of object categories absent from existing 3D datasets. Experiments show that annotations generated by LabelAny3D improve monocular 3D detection performance across multiple benchmarks, outperforming prior auto-labeling approaches in quality. These results demonstrate the promise of foundation-model-driven annotation for scaling up 3D recognition in realistic, open-world settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
相关 Paper
- Omni3D: A Large Benchmark and Model for 3D Object Detection in the WildGarrick Brazil, Abhinav Kumar, Julian Straub, Nikhila Ravi 等CVPR 2023
- MonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular DetectionRishubh Parihar, Srinjay Sarkar, Sarthak Vora, Jogendra Nath Kundu 等CVPR 2025
- MEBOW: Monocular Estimation of Body Orientation in the WildChenyan Wu, Yukun Chen, Jiajia Luo, Che-Chun Su 等CVPR 2020
- OpenBox: Annotate Any Bounding Boxes in 3DIn-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio 等NeurIPS 2025 · 被引用 7 次
- Learning Occupancy for Monocular 3D Object DetectionLiang Peng, Junkai Xu, Haoran Cheng, Zheng Yang 等CVPR 2024 · 被引用 21 次
