SIR: Structured Image Representations for Explainable Robot Learning
Paul Mattes, Jan Schwab, Jens Bosch, Maximilian Xiling Li, Nils Blank, Minh-Trung Tang, Moritz Haberland, Rudolf Lioutikov
Abstract
Existing robot policies based on learned visual embeddings lack explicit structure and are sensitive to visual distractions. Thus, the representations that drive their behaviour are often opaque, making their decision-making process difficult to interpret. To address this, we introduce Structured Image Representations (SIR), a method that leverages Scene Graphs (SGs) as an intermediate representation for robot policy learning. Our approach first constructs a fully connected graph, using image-derived features as initial node representations. Then, a module learns to sparsify this graph end-to-end, creating a task-relevant sub-graph that is passed to the action generation model. This process makes our model intrinsically explainable. Evaluations on RoboCasa show that our sparse graph policies outperform image-based baselines on average with 19.5% vs 14.81% success rate. Most importantly, we show that the learned sparse graphs are a powerful tool for model analysis. By analysing when the model's sub-graph deviates from human expectation, such as by including distractor nodes or omitting key objects, we successfully uncover dataset biases, including spurious correlations and positional biases. https://github.com/intuitive-robots/SIR_Model
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9254fdb0-6505-4ace-afa0-01b3cf8a28d1Builds on9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Maximum Likelihood Training of Score-Based Diffusion ModelsYang Song, Conor Durkan, Iain Murray, Stefano ErmonNeurIPS 2021 · 958 citations
- Behavior Transformers: Cloning modes with one stoneNur Muhammad Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, Lerrel PintoNeurIPS 2022 · 470 citations
- Robust Graph Representation Learning via Neural SparsificationCheng Zheng, Bo Zong, Wei Cheng, Dongjin Song et al.ICML 2020 · 330 citations
- Scene Graph Contrastive Learning for Embodied NavigationKunal Pratap Singh, Jordi Salvador, Luca Weihs, Aniruddha KembhaviICCV 2023 · 31 citations
Related papers
- Expanding Spatial and Temporal Context for Robotic Imitation Learning With Scene GraphsJianing Qian, Qinhe Peng, Emmanuel Panov, Leonor Fermoselle et al.CVPR 2026
- Scene Graph-Grounded Image GenerationFuyun Wang, Tong Zhang, Yuanzhi Wang, Xiaoya Zhang et al.AAAI 2025 · 1 citation
- Structured Sparse R-CNN for Direct Scene Graph GenerationYao Teng, Limin WangCVPR 2022 · 66 citations
- CURVE: Learning Causality-Inspired Invariant Representations for Robust Scene Understanding via Uncertainty-Guided RegularizationYue Liang, JIATONG DU, Ziyi Yang, Yanjun Huang et al.ICML 2026 · 2 citations
- Generating Explanations for Embodied Action Decision from Visual ObservationXiaohan Wang, Yuehu Liu, Xinhang Song, Beibei Wang et al.ACM MM 2023 · 3 citations
