PhoCaL: A Multi-Modal Dataset for Category-Level Object Pose Estimation with Photometrically Challenging Objects
Pengyuan Wang, HyunJun Jung, Yitong Li, Siyuan Shen, Rahul Parthasarathy Srikanth, Lorenzo Garattoni, Sven Meier, Nassir Navab, Benjamin Busam
Abstract
Object pose estimation is crucial for robotic applications and augmented reality. Beyond instance level 6D object pose estimation methods, estimating category-level pose and shape has become a promising trend. As such, a new research field needs to be supported by well-designed datasets. To provide a benchmark with high-quality ground truth annotations to the community, we introduce a multimodal dataset for category-level object pose estimation with photometrically challenging objects termed PhoCaL. PhoCaL comprises 60 high quality 3D models of household objects over 8 categories including highly reflective, transparent and symmetric objects. We developed a novel robot-supported multi-modal (RGB, depth, polarisation) data acquisition and annotation process. It ensures sub-millimeter accuracy of the pose for opaque textured, shiny and transparent objects, no motion blur and perfect camera synchronisation. To set a benchmark for our dataset, state-of-the-art RGB-D and monocular RGB methods are evaluated on the challenging scenes of PhoCaL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 04fbbb74-a435-4e64-9211-1f3ebbe0c559Cited by top-tier papers16
- RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D VideosHongchi Xia, Yang Fu, Sifei Liu, Xiaolong WangCVPR 2024 · 14 citations
- CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge DistillationXiao Lin, Yun Peng, Liuyi Wang, Xianyou Zhong et al.ICCV 2025 · 3 citations
- MultiCam: On-the-fly Multi-Camera Pose Estimation Using Spatiotemporal Overlaps of Known ObjectsShiyu Li, Hannah Schieber, Kristoffer Waldow, Benjamin Busam et al.IEEE VR 2026 · 2 citations
- AutoSynth: Learning to Generate 3D Training Data for Object Point Cloud RegistrationZheng Dang, Mathieu SalzmannICCV 2023 · 1 citation
- Seeing and Seeing Through the Glass: Real and Synthetic Data for Multi-Layer Depth EstimationHongyu Wen, Yiming Zuo, Venkat Subramanian, Patrick Chen et al.ICCV 2025 · 1 citation
Builds on9
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 486 citations
- SGPA: Structure-Guided Prior Adaptation for Category-Level 6D Object Pose EstimationKai Chen, Qi DouICCV 2021 · 183 citations
- DualPoseNet: Category-level 6D Object Pose and Size Estimation Using Dual Pose Network with Refined Learning of Pose ConsistencyJiehong Lin, Zewei Wei, Zhihao Li, Songcen Xu et al.ICCV 2021 · 169 citations
- StereOBJ-1M: Large-scale Stereo Image Dataset for 6D Object Pose EstimationXingyu Liu, Shun Iwase, Kris M. KitaniICCV 2021 · 58 citations
- GraspNet-1Billion: A Large-Scale Benchmark for General Object GraspingHaoshu Fang, Chenxi Wang, Minghao Gou, Cewu LuCVPR 2020
Related papers
- HouseCat6D - A Large-Scale Multi-Modal Category Level 6D Object Perception Dataset with Household Objects in Realistic ScenariosHyunJun Jung, Shun-Cheng Wu, Patrick Ruhkamp, Guangyao Zhai et al.CVPR 2024
- Towards Co-Evaluation of Cameras, HDR, and Algorithms for Industrial-Grade 6DoF Pose EstimationAgastya Kalra, Guy Stoppi, Dmitrii Marin, Vage Taamazyan et al.CVPR 2024
- Towards Multimodal Depth Estimation from Light FieldsTitus Leistner, Radek Mackowiak, Lynton Ardizzone, Ullrich Köthe et al.CVPR 2022 · 14 citations
- GP2C: Geometric Projection Parameter Consensus for Joint 3D Pose and Focal Length Estimation in the WildAlexander Grabner, Peter M. Roth, Vincent LepetitICCV 2019 · 21 citations
- PolarDepth: Monocular Transparent Object Depth from Polar-Physics PriorsWen Dong, Haiyang Mei, Yinglian Ji, Zijun Zhang et al.ICML 2026
