Memory-based Adapters for Online 3D Scene Perception
Xiuwei Xu, Chong Xia, Ziwei Wang, Linqing Zhao, Yueqi Duan, Jie Zhou, Jiwen Lu
摘要
In this paper, we propose a new framework for online 3D scene perception. Conventional 3D scene perception methods are offline, i.e., take an already reconstructed 3D scene geometry as input, which is not applicable in robotic applications where the input data is streaming RGB-D videos rather than a complete 3D scene reconstructed from precollected RGB-D videos. To deal with online 3D scene perception tasks where data collection and perception should be performed simultaneously, the model should be able to process 3D scenes frame by frame and make use of the temporal information. To this end, we propose an adapter-based plug-and-play module for the backbone of 3D scene perception model, which constructs memory to cache and aggregate the extracted RGB-D features to empower offline models with temporal learning ability. Specifically, we propose a queued memory mechanism to cache the supporting point cloud and image features. Then we devise aggregation modules which directly perform on the memory and pass temporal information to current frame. We further propose 3D-to-2D adapter to enhance image features with strong global context. Our adapters can be easily inserted into mainstream offline architectures of different tasks and significantly boost their performance on online tasks. Extensive experiments on ScanNet and SceneNN datasets demonstrate our approach achieves leading performance on three 3D scene perception tasks compared with state-of-the-art online methods by simply finetuning existing offline models, without any model and task-specific designs. Project page.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object NavigationHang Yin, Xiuwei Xu, Zhenyu Wu, Jie Zhou 等NeurIPS 2024 · 被引用 215 次
- LowRankOcc: Tensor Decomposition and Low-Rank Recovery for Vision-Based 3D Semantic Occupancy PredictionLinqing Zhao, Xiuwei Xu, Ziwei Wang, Yunpeng Zhang 等CVPR 2024 · 被引用 14 次
- Online Segment Any 3D Thing as Instance TrackingHanshi Wang, Zijian Cai, Jin Gao, Yiwei Zhang 等NeurIPS 2025 · 被引用 6 次
- EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-Based Online Scene UnderstandingYuqi Wu, Wenzhao Zheng, Sicheng Zuo, Yuanhui Huang 等ICCV 2025 · 被引用 4 次
- EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene UnderstandingSeungjun Lee, Zihan Wang, Yunsong Wang, Gim Hee LeeCVPR 2026 · 被引用 2 次
它引用的顶会 Paper16
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 被引用 857 次
- 6-DOF GraspNet: Variational Grasp Generation for Object ManipulationArsalan Mousavian, Clemens Eppner, Dieter FoxICCV 2019 · 被引用 673 次
- ST-Adapter: Parameter-Efficient Image-to-Video Transfer LearningJunting Pan, Ziyi Lin, Xiatian Zhu, Jing Shao 等NeurIPS 2022 · 被引用 290 次
相关 Paper
- Continuous 3D Perception Model with Persistent StateQianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A. Efros 等CVPR 2025
- Real-time Stereo-based 3D Object Detection for Streaming PerceptionChangcai Li, Zonghua Gu, Gang Chen, Libo Huang 等NeurIPS 2024 · 被引用 3 次
- ESAM++: Efficient Online 3D Perception on the EdgeQin Liu, Lavisha Aggarwal, Saptarashmi Bandyopadhyay, Vikas Bahirwani 等CVPR 2026 · 被引用 1 次
- EmbodiedSAM: Online Segment Any 3D Thing in Real TimeXiuwei Xu, Huangxing Chen, Linqing Zhao, Ziwei Wang 等ICLR 2025
- Occupancy Learning with Spatiotemporal MemoryZiyang Leng, Jiawei Yang, Wenlong Yi, Bolei ZhouICCV 2025 · 被引用 10 次
