3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
Yuncong Yang, Han Yang, Jiachen Zhou, Peihao Chen, Hongxin Zhang, Yilun Du, Chuang Gan
2025Year
17Top-tier citations
Abstract
Figure 1. With 3D-Mem, explored regions are represented by a set of Memory Snapshots capturing clusters of co-visible objects, i.e., the objects observable in a single image observation, along with their spatial relationships and background context, as shown in the bottom-left example. Unexplored regions are represented by navigable frontiers along with image observations, referred to as Frontier Snapshots.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a52af42-525e-479e-8104-deab9dc8a691Cited by top-tier papers17
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual TokensZeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen et al.CVPR 2026 · 124 citations
- Learning 3D Persistent Embodied World ModelsSiyuan Zhou, Yilun Du, Yuncong Yang, Lei Han et al.NeurIPS 2025 · 34 citations
- Describe Anything Anywhere At Any MomentNicolas Gorlo, Lukas Schmid, Luca CarloneCVPR 2026 · 27 citations
- MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied NavigationXun Huang, Shijia Zhao, Yunxiang Wang, Xin Lu et al.CVPR 2026 · 19 citations
- SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and HearingMingfei Chen, Zijun Cui, Xiulong Liu, Jinlin Xiang et al.NeurIPS 2025 · 18 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
Related papers
- VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View MemoryRunjia Li, Philip Torr, Andrea Vedaldi, Tomas JakabICCV 2025 · 11 citations
- CogniMap3D: Cognitive 3D Mapping and Rapid RetrievalFeiran Wang, Junyi Wu, Dawen Cai, Yuan Hong et al.ICLR 2026 · 1 citation
- What You See is What You Get: Exploiting Visibility for 3D Object DetectionPeiyun Hu, Jason Ziglar, David Held, Deva RamananCVPR 2020
- Memory-Augmented Scene Understanding and Exploration for Open-World Aerial Object-Goal NavigationJiacong Zhou, Jiaxu Miao, Yourun Lin, Xianyun Wang et al.CVPR 2026
- SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object NavigationHang Yin, Xiuwei Xu, Zhenyu Wu, Jie Zhou et al.NeurIPS 2024 · 215 citations
