RIO: 3D Object Instance Re-Localization in Changing Indoor Environments
Johanna Wald, Armen Avetisyan, Nassir Navab, Federico Tombari, Matthias Nießner
Abstract
In this work, we introduce the task of 3D object instance re-localization (RIO): given one or multiple objects in an RGB-D scan, we want to estimate their corresponding 6DoF poses in another 3D scan of the same environment taken at a later point in time. We consider RIO a particularly important task in 3D vision since it enables a wide range of practical applications, including AI-assistants or robots that are asked to find a specific object in a 3D scene. To address this problem, we first introduce 3RScan, a novel dataset and benchmark, which features 1482 RGB-D scans of 478 environments across multiple time steps. Each scene includes several objects whose positions change over time, together with ground truth annotations of object instances and their respective 6DoF mappings among re-scans. Automatically finding 6DoF object poses leads to a particular challenging feature matching task due to varying partial observations and changes in the surrounding context. To this end, we introduce a new data-driven approach that efficiently finds matching features using a fully-convolutional 3D correspondence network operating on multiple spatial scales. Combined with a 6DoF pose optimization, our method outperforms state-of-the-art baselines on our newly-established benchmark, achieving an accuracy of 30.58%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32c3fb76-68cc-45b1-a463-03810ff9bdb3Cited by top-tier papers86
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
- 3D-VisTA: Pre-trained Transformer for 3D Vision and Text AlignmentZiyu Zhu, Xiaojian Ma, Yixin Chen, Zhidong Deng et al.ICCV 2023 · 247 citations
- GPT4Scene: Understand 3D Scenes from Videos with Vision-Language ModelsZhangyang Qi, Zhixiong Zhang, Ye Fang, Jiaqi Wang et al.ICLR 2026 · 121 citations
- Graph-to-3D: End-to-End Generation and Manipulation of 3D Scenes Using Scene GraphsHelisa Dhamo, Fabian Manhardt, Nassir Navab, Federico TombariICCV 2021 · 98 citations
- SpatialLM: Training Large Language Models for Structured Indoor ModelingYongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng et al.NeurIPS 2025 · 89 citations
Related papers
- 3D Focusing-and-Matching Network for Multi-Instance Point Cloud RegistrationLiyuan Zhang, Le Hui, Qi Liu, Bo Li et al.NeurIPS 2024 · 4 citations
- ReScene4D: Temporally Consistent Semantic Instance Segmentation of Evolving Indoor 3D ScenesEmily Steiner, Jianhao Zheng, Henry Howard-Jenkins, Chris Xie et al.CVPR 2026 · 3 citations
- Living Scenes: Multi-object Relocalization and Reconstruction in Changing 3D EnvironmentsLiyuan Zhu, Shengyu Huang, Konrad Schindler, Iro ArmeniCVPR 2024 · 10 citations
- FroDO: From Detections to 3D ObjectsMartin Rünz, Kejie Li, Meng Tang, Lingni Ma et al.CVPR 2020
- RevealNet: Seeing Behind Objects in RGB-D ScansJi Hou, Angela Dai, Matthias NießnerCVPR 2020
