StereOBJ-1M: Large-scale Stereo Image Dataset for 6D Object Pose Estimation
Xingyu Liu, Shun Iwase, Kris M. Kitani
Abstract
We present a large-scale stereo RGB image object pose estimation dataset named the StereOBJ-1M dataset. The dataset is designed to address challenging cases such as object transparency, translucency, and specular reflection, in addition to the common challenges of occlusion, symmetry, and variations in illumination and environments. In order to collect data of sufficient scale for modern deep learning models, we propose a novel method for efficiently annotating pose data in a multi-view fashion that allows data capturing in complex and flexible environments. Fully annotated with 6D object poses, our dataset contains over 393K frames and over 1.5M annotations of 18 objects recorded in 182 scenes constructed in 11 different environments. The 18 objects include 8 symmetric objects, 7 transparent objects, and 8 reflective objects. We benchmark two state-of-the-art pose estimation frameworks on StereOBJ-1M as baselines for future work. We also propose a novel object-level pose optimization method for computing 6D pose from keypoint predictions in multiple images. Project website: https: //sites.google.com/view/stereobj-1m .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e1136d8-5cc8-45bf-af5e-0fad09357a92Cited by top-tier papers11
- PhoCaL: A Multi-Modal Dataset for Category-Level Object Pose Estimation with Photometrically Challenging ObjectsPengyuan Wang, HyunJun Jung, Yitong Li, Siyuan Shen et al.CVPR 2022 · 44 citations
- Learning Depth Estimation for Transparent and Mirror SurfacesAlex Costanzino, Pierluigi Zama Ramirez, Matteo Poggi, Fabio Tosi et al.ICCV 2023 · 41 citations
- Stability-driven Contact Reconstruction From Monocular Color ImagesZimeng Zhao, Binghui Zuo, Wei Xie, Yangang WangCVPR 2022 · 15 citations
- BCOT: A Markerless High-Precision 3D Object Tracking BenchmarkJiachen Li, Bin Wang, Shiqiang Zhu, Xin Cao et al.CVPR 2022 · 15 citations
- RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D VideosHongchi Xia, Yang Fu, Sifei Liu, Xiaolong WangCVPR 2024 · 14 citations
Builds on5
- Point-Based Multi-View Stereo NetworkRui Chen, Songfang Han, Jing Xu, Hao SuICCV 2019 · 403 citations
- MeteorNet: Deep Learning on Dynamic 3D Point Cloud SequencesXingyu Liu, Mengyuan Yan, Jeannette BohgICCV 2019 · 225 citations
- PVN3D: A Deep Point-Wise 3D Keypoints Voting Network for 6DoF Pose EstimationYisheng He, Wei Sun, Haibin Huang, Jianran Liu et al.CVPR 2020
- GraspNet-1Billion: A Large-Scale Benchmark for General Object GraspingHaoshu Fang, Chenxi Wang, Minghao Gou, Cewu LuCVPR 2020
- KeyPose: Multi-View 3D Labeling and Keypoint Estimation for Transparent ObjectsXingyu Liu, Rico Jonschkowski, Anelia Angelova, Kurt KonoligeCVPR 2020
Related papers
- Open Challenges in Deep Stereo: the Booster DatasetPierluigi Zama Ramirez, Fabio Tosi, Matteo Poggi, Samuele Salti et al.CVPR 2022 · 39 citations
- KeypointNet: A Large-Scale 3D Keypoint Dataset Aggregated From Numerous Human AnnotationsYang You, Yujing Lou, Chengkun Li, Zhoujun Cheng et al.CVPR 2020
- Towards Multimodal Depth Estimation from Light FieldsTitus Leistner, Radek Mackowiak, Lynton Ardizzone, Ullrich Köthe et al.CVPR 2022 · 14 citations
- DiLiGenT102: A Photometric Stereo Benchmark Dataset with Controlled Shape and Material VariationJieji Ren, Feishi Wang, Jiahao Zhang, Qian Zheng et al.CVPR 2022 · 26 citations
- 3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture ObjectsZhicheng Liang, Haoyi Yu, Boyan Li, Dayou Zhang et al.CVPR 2026 · 1 citation
