C-Reference: Improving 2D to 3D Object Pose Estimation Accuracy via Crowdsourced Joint Object Estimation
Jean Y. Song, John Joon Young Chung, David F. Fouhey, Walter S. Lasecki
Abstract
Converting widely-available 2D images and videos, captured using an RGB camera, to 3D can help accelerate the training of machine learning systems in spatial reasoning domains ranging from in-home assistive robots to augmented reality to autonomous vehicles. However, automating this task is challenging because it requires not only accurately estimating object location and orientation, but also requires knowing currently unknown camera properties (e.g., focal length). A scalable way to combat this problem is to leverage people's spatial understanding of scenes by crowdsourcing visual annotations of 3D object properties. Unfortunately, getting people to directly estimate 3D properties reliably is difficult due to the limitations of image resolution, human motor accuracy, and people's 3D perception (i.e., humans do not "see" depth like a laser range finder). In this paper, we propose a crowd-machine hybrid approach that jointly uses crowds' approximate measurements of multiple in-scene objects to estimate the 3D state of a single target object. Our approach can generate accurate estimates of the target object by combining heterogeneous knowledge from multiple contributors regarding various different objects that share a spatial relationship with the target object. We evaluate our joint object estimation approach with 363 crowd workers and show that our method can reduce errors in the target object's 3D location estimation by over 40%, while requiring only % as much human time. Our work introduces a novel way to enable groups of people with different perspectives and knowledge to achieve more accurate collective performance on challenging visual annotation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Crowdsourcing More Effective Initializations for Single-Target Trackers Through Automatic Re-queryingStephan J. Lemmer, Jean Y. Song, Jason J. CorsoCHI 2021 · 6 citations
- Ground-truth or DAER: Selective Re-query of Secondary InformationStephan J. Lemmer, Jason J. CorsoICCV 2021 · 4 citations
Builds on2
- 3D-RelNet: Joint Object and Relational Network for 3D PredictionNilesh Kulkarni, Ishan Misra, Shubham Tulsiani, Abhinav GuptaICCV 2019 · 48 citations
- Improving Crowd-Supported GUI Testing with Structural GuidanceYan Chen, Maulishree Pandey, Jean Y. Song, Walter S. Lasecki et al.CHI 2020 · 26 citations
Related papers
- CAD-Estate: Large-scale CAD Model Annotation in RGB VideosKevis-Kokitsi Maninis, Stefan Popov, Matthias Nießner, Vittorio FerrariICCV 2023 · 14 citations
- CrowdDriven: A New Challenging Dataset for Outdoor Visual LocalizationAra Jafarzadeh, Manuel López-Antequera, Pau Gargallo, Yubin Kuang et al.ICCV 2021 · 18 citations
- CHORUS: Learning Canonicalized 3D Human-Object Spatial Relations from Unbounded Synthesized ImagesSookwan Han, Hanbyul JooICCV 2023 · 19 citations
- MonoSOWA: Scalable Monocular 3D Object Detector Without Human AnnotationsJan Skvrna, Lukás NeumannICCV 2025 · 3 citations
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri et al.CVPR 2025
