SIR: Structured Image Representations for Explainable Robot Learning
Paul Mattes, Jan Schwab, Jens Bosch, Maximilian Xiling Li, Nils Blank, Minh-Trung Tang, Moritz Haberland, Rudolf Lioutikov
摘要
Existing robot policies based on learned visual embeddings lack explicit structure and are sensitive to visual distractions. Thus, the representations that drive their behaviour are often opaque, making their decision-making process difficult to interpret. To address this, we introduce Structured Image Representations (SIR), a method that leverages Scene Graphs (SGs) as an intermediate representation for robot policy learning. Our approach first constructs a fully connected graph, using image-derived features as initial node representations. Then, a module learns to sparsify this graph end-to-end, creating a task-relevant sub-graph that is passed to the action generation model. This process makes our model intrinsically explainable. Evaluations on RoboCasa show that our sparse graph policies outperform image-based baselines on average with 19.5% vs 14.81% success rate. Most importantly, we show that the learned sparse graphs are a powerful tool for model analysis. By analysing when the model's sub-graph deviates from human expectation, such as by including distractor nodes or omitting key objects, we successfully uncover dataset biases, including spurious correlations and positional biases. https://github.com/intuitive-robots/SIR_Model
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Maximum Likelihood Training of Score-Based Diffusion ModelsYang Song, Conor Durkan, Iain Murray, Stefano ErmonNeurIPS 2021 · 被引用 958 次
- Behavior Transformers: Cloning modes with one stoneNur Muhammad Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, Lerrel PintoNeurIPS 2022 · 被引用 470 次
- Robust Graph Representation Learning via Neural SparsificationCheng Zheng, Bo Zong, Wei Cheng, Dongjin Song 等ICML 2020 · 被引用 330 次
- Scene Graph Contrastive Learning for Embodied NavigationKunal Pratap Singh, Jordi Salvador, Luca Weihs, Aniruddha KembhaviICCV 2023 · 被引用 31 次
相关 Paper
- Expanding Spatial and Temporal Context for Robotic Imitation Learning With Scene GraphsJianing Qian, Qinhe Peng, Emmanuel Panov, Leonor Fermoselle 等CVPR 2026
- Scene Graph-Grounded Image GenerationFuyun Wang, Tong Zhang, Yuanzhi Wang, Xiaoya Zhang 等AAAI 2025 · 被引用 1 次
- Structured Sparse R-CNN for Direct Scene Graph GenerationYao Teng, Limin WangCVPR 2022 · 被引用 66 次
- CURVE: Learning Causality-Inspired Invariant Representations for Robust Scene Understanding via Uncertainty-Guided RegularizationYue Liang, JIATONG DU, Ziyi Yang, Yanjun Huang 等ICML 2026 · 被引用 2 次
- Generating Explanations for Embodied Action Decision from Visual ObservationXiaohan Wang, Yuehu Liu, Xinhang Song, Beibei Wang 等ACM MM 2023 · 被引用 3 次
