Memorize What Matters: Emergent Scene Decomposition from Multitraverse
Yiming Li, Zehong Wang, Yue Wang, Zhiding Yu, Zan Gojcic, Marco Pavone, Chen Feng, José M. Álvarez
Abstract
Humans naturally retain memories of permanent elements, while ephemeral moments often slip through the cracks of memory. This selective retention is crucial for robotic perception, localization, and mapping. To endow robots with this capability, we introduce 3D Gaussian Mapping (3DGM), a self-supervised, camera-only offline mapping framework grounded in 3D Gaussian Splatting. 3DGM converts multitraverse RGB videos from the same region into a Gaussian-based environmental map while concurrently performing 2D ephemeral object segmentation. Our key observation is that the environment remains consistent across traversals, while objects frequently change. This allows us to exploit self-supervision from repeated traversals to achieve environment-object decomposition. More specifically, 3DGM formulates multitraverse environmental mapping as a robust differentiable rendering problem, treating pixels of the environment and objects as inliers and outliers, respectively. Using robust feature distillation, feature residuals mining, and robust optimization, 3DGM jointly performs 2D segmentation and 3D mapping without human intervention. We build the Mapverse benchmark, sourced from the Ithaca365 and nuPlan datasets, to evaluate our method in unsupervised 2D segmentation, 3D reconstruction, and neural rendering. Extensive results verify the effectiveness and potential of our method for self-driving and robotics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5f777b1-8ac4-4e23-8b64-db7687a4a5ccCited by top-tier papers3
- REDOUBT: Duo Safety Validation for Autonomous Vehicle Motion PlanningShuguang Wang, Qian Zhou, Kui Wu, Dapeng Wu et al.NeurIPS 2025 · 6 citations
- Adversarial Exploitation of Data Diversity Improves Visual LocalizationSihang Li, Siqi Tan, Bowen Chang, Jing Zhang et al.ICCV 2025 · 4 citations
- Extrapolated Urban View Synthesis BenchmarkXiangyu Han, Zhen Jia, Boyi Li, Yan Wang et al.ICCV 2025 · 3 citations
Builds on50
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- DeGauss: Dynamic-Static Decomposition with Gaussian Splatting for Distractor-Free 3D ReconstructionRui Wang, Quentin Lohmeyer, Mirko Meboldt, Siyu TangICCV 2025 · 12 citations
- Gaussian Mapping for Evolving ScenesVladimir Yugay, Thies Kersten, Luca Carlone, Theo Gevers et al.CVPR 2026 · 5 citations
- OnlinePG: Online Open-Vocabulary Panoptic Mapping with 3D Gaussian SplattingHongjia Zhai, Qi Zhang, Xiaokun Pan, Xiyu Zhang et al.CVPR 2026 · 3 citations
- DG-SLAM: Robust Dynamic Gaussian Splatting SLAM with Hybrid Pose OptimizationYueming Xu, Haochen Jiang, Zhongyang Xiao, Jianfeng Feng et al.NeurIPS 2024 · 65 citations
- Video2Robo: 3DGS-based Synthetic Data from One Video Enables Scalable Robot LearningYinan Deng, Kejia Hu, Ye Chen, Jianyu Dou et al.CVPR 2026
