RIO: 3D Object Instance Re-Localization in Changing Indoor Environments
Johanna Wald, Armen Avetisyan, Nassir Navab, Federico Tombari, Matthias Nießner
摘要
In this work, we introduce the task of 3D object instance re-localization (RIO): given one or multiple objects in an RGB-D scan, we want to estimate their corresponding 6DoF poses in another 3D scan of the same environment taken at a later point in time. We consider RIO a particularly important task in 3D vision since it enables a wide range of practical applications, including AI-assistants or robots that are asked to find a specific object in a 3D scene. To address this problem, we first introduce 3RScan, a novel dataset and benchmark, which features 1482 RGB-D scans of 478 environments across multiple time steps. Each scene includes several objects whose positions change over time, together with ground truth annotations of object instances and their respective 6DoF mappings among re-scans. Automatically finding 6DoF object poses leads to a particular challenging feature matching task due to varying partial observations and changes in the surrounding context. To this end, we introduce a new data-driven approach that efficiently finds matching features using a fully-convolutional 3D correspondence network operating on multiple spatial scales. Combined with a 6DoF pose optimization, our method outperforms state-of-the-art baselines on our newly-established benchmark, achieving an accuracy of 30.58%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper86
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu 等ICML 2024 · 被引用 361 次
- 3D-VisTA: Pre-trained Transformer for 3D Vision and Text AlignmentZiyu Zhu, Xiaojian Ma, Yixin Chen, Zhidong Deng 等ICCV 2023 · 被引用 247 次
- GPT4Scene: Understand 3D Scenes from Videos with Vision-Language ModelsZhangyang Qi, Zhixiong Zhang, Ye Fang, Jiaqi Wang 等ICLR 2026 · 被引用 121 次
- Graph-to-3D: End-to-End Generation and Manipulation of 3D Scenes Using Scene GraphsHelisa Dhamo, Fabian Manhardt, Nassir Navab, Federico TombariICCV 2021 · 被引用 98 次
- SpatialLM: Training Large Language Models for Structured Indoor ModelingYongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng 等NeurIPS 2025 · 被引用 89 次
相关 Paper
- 3D Focusing-and-Matching Network for Multi-Instance Point Cloud RegistrationLiyuan Zhang, Le Hui, Qi Liu, Bo Li 等NeurIPS 2024 · 被引用 4 次
- ReScene4D: Temporally Consistent Semantic Instance Segmentation of Evolving Indoor 3D ScenesEmily Steiner, Jianhao Zheng, Henry Howard-Jenkins, Chris Xie 等CVPR 2026 · 被引用 3 次
- Living Scenes: Multi-object Relocalization and Reconstruction in Changing 3D EnvironmentsLiyuan Zhu, Shengyu Huang, Konrad Schindler, Iro ArmeniCVPR 2024 · 被引用 10 次
- FroDO: From Detections to 3D ObjectsMartin Rünz, Kejie Li, Meng Tang, Lingni Ma 等CVPR 2020
- RevealNet: Seeing Behind Objects in RGB-D ScansJi Hou, Angela Dai, Matthias NießnerCVPR 2020
