Modeling Dynamic Environments with Scene Graph Memory
Andrey Kurenkov, Michael Lingelbach, Tanmay Agarwal, Emily Jin, Chengshu Li, Ruohan Zhang, Li Fei-Fei, Jiajun Wu, Silvio Savarese, Roberto Martín-Martín
Abstract
Embodied AI agents that search for objects in large environments such as households often need to make efficient decisions by predicting object locations based on partial information. We pose this as a new type of link prediction problem: link prediction on partially observable dynamic graphs. Our graph is a representation of a scene in which rooms and objects are nodes, and their relationships are encoded in the edges; only parts of the changing graph are known to the agent at each timestep. This partial observability poses a challenge to existing link prediction approaches, which we address. We propose a novel state representation -Scene Graph Memory (SGM) -with captures the agent's accumulated set of observations, as well as a neural net architecture called a Node Edge Predictor (NEP) that extracts information from the SGM to search efficiently. We evaluate our method in the Dynamic House Simulator, a new benchmark that creates diverse dynamic graphs following the semantic patterns typically seen at homes, and show that NEP can be trained to predict the locations of objects in a variety of environments with diverse object movement dynamics, outperforming baselines both in terms of new scene adaptability and overall accuracy. The codebase and more can be found this URL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Align While Search: Belief-Guided Exploratory Inference for World-Grounded Embodied AgentsSeohui Bae, Jeonghye Kim, Youngchul Sung, Woohyung LimCVPR 2026 · 1 citation
- Merge3D: Efficient 3D Multimodal LLMs via Joint 2D-3D Token MergingTianbo Pan, Xingyi Yang, Xinchao WangCVPR 2026
- Zero-shot 3D Question Answering via Voxel-based Dynamic Token CompressionHsiang-Wei Huang, Fu-Chen Chen, Wenhao Chai, Che-Chun Su et al.CVPR 2025
- Adaptive Memory Retention in Dynamic GraphsFabrizio De Castelli, Alessio Gravina, Moshe Eliasof, Carola-Bibiane Schönlieb et al.ICML 2026
Builds on10
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- 🏘️ ProcTHOR: Large-Scale Embodied AI Using Procedural GenerationMatt Deitke, Eli VanderBilt, Alvaro Herrasti, Luca Weihs et al.NeurIPS 2022 · 596 citations
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli et al.ICLR 2020 · 584 citations
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak et al.ICCV 2019 · 474 citations
- GraphFormers: GNN-nested Transformers for Representation Learning on Textual GraphJunhan Yang, Zheng Liu, Shitao Xiao, Chaozhuo Li et al.NeurIPS 2021 · 262 citations
Related papers
- Expanding Spatial and Temporal Context for Robotic Imitation Learning With Scene GraphsJianing Qian, Qinhe Peng, Emmanuel Panov, Leonor Fermoselle et al.CVPR 2026
- Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene UnderstandingYue Fan, Xiaojian Ma, Rongpeng Su, Jun Guo et al.ICCV 2025 · 2 citations
- Scene Graph Contrastive Learning for Embodied NavigationKunal Pratap Singh, Jordi Salvador, Luca Weihs, Aniruddha KembhaviICCV 2023 · 31 citations
- Continuous Scene Representations for Embodied AISamir Yitzhak Gadre, Kiana Ehsani, Shuran Song, Roozbeh MottaghiCVPR 2022 · 40 citations
- Semi-Weakly Supervised Object Kinematic Motion PredictionGengxin Liu, Qian Sun, Haibin Huang, Chongyang Ma et al.CVPR 2023
