Expanding Spatial and Temporal Context for Robotic Imitation Learning With Scene Graphs
Jianing Qian, Qinhe Peng, Emmanuel Panov, Leonor Fermoselle, Dinesh Jayaraman, Bernadette Bucher, Tarik Kelestemur
Abstract
Imitation learning enables robots to learn how to execute tasks via observation. However, real-world environments like homes and offices are often severely partially observed due to their large spatial scales. In addition, many tasks involve executing a series of subtasks requiring autonomous robots to reason over extended time horizons. To address these challenges, we propose using scene graphs as an explicit and structured memory mechanism in imitation learning. By maintaining a dynamic scene graph that captures object-centric relationships and their evolution over time, our method allows the agent to retain relevant historical context during task execution to efficiently reason over incrementally accrued scene information. Our experiments on simulated mobile manipulation and real-world tabletop manipulation demonstrate that our approach substantially improves policy performance, particularly in settings that demand long-term reasoning and robust generalization under partial observability. Code and videos:https://sites.google.com/view/objgraph.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69867cbb-2db4-4689-95d0-17c729f637e6Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- 3D Scene Graph: A Structure for Unified Semantics, 3D Space, and CameraIro Armeni, Zhi-Yang He, Amir Zamir, JunYoung Gwak et al.ICCV 2019 · 474 citations
- Knowledge-inspired 3D Scene Graph Prediction in Point CloudShoulong Zhang, Shuai Li, Aimin Hao, Hong QinNeurIPS 2021 · 54 citations
- ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon TasksKaijun Wang, Liqin Lu, Mingyu Liu, Jianuo Jiang et al.AAAI 2026 · 6 citations
Related papers
- Modeling Dynamic Environments with Scene Graph MemoryAndrey Kurenkov, Michael Lingelbach, Tanmay Agarwal, Emily Jin et al.ICML 2023 · 21 citations
- SIR: Structured Image Representations for Explainable Robot LearningPaul Mattes, Jan Schwab, Jens Bosch, Maximilian Xiling Li et al.CVPR 2026 · 1 citation
- Deep Imitation Learning for Bimanual Robotic ManipulationFan Xie, Alexander Chowdhury, M. Clara De Paolis Kaluza, Linfeng Zhao et al.NeurIPS 2020 · 108 citations
- MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Models for Embodied Task PlanningYuanchen Ju, Yongyuan Liang, Yen-Jen Wang, Nandiraju Gireesh et al.ICLR 2026 · 5 citations
- Hierarchical 3D Scene Graphs Construction OutdoorsJon Nyffeler, Federico Tombari, Daniel BarathICCV 2025 · 1 citation
