ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
Jiani Huang, Amish Sethi, Matthew Kuo, Mayank Keoliya, Neelay Velingker, JungHo Jung, Ser Nam Lim, Ziyang Li, Mayur Naik
Abstract
Multi-modal large language models (MLLMs) are making rapid progress toward general-purpose embodied agents. However, existing MLLMs do not reliably capture fine-grained links between low-level visual features and high-level textual semantics, leading to weak grounding and inaccurate perception. To overcome this challenge, we propose ESCA, a framework that contextualizes embodied agents by grounding their perception in spatial-temporal scene graphs. At its core is SGClip, a novel, open-domain, promptable foundation model for generating scene graphs that is based on CLIP. SGClip is trained on 87K+ open-domain videos using a neurosymbolic pipeline that aligns automatically generated captions with scene graphs produced by the model itself, eliminating the need for human-labeled annotations. We demonstrate that SGClip excels in both prompt-based inference and task-specific fine-tuning, achieving state-of-the-art results on scene graph generation and action localization benchmarks. ESCA with SGClip improves perception for embodied agents based on both open-source and commercial MLLMs, achieving state of-the-art performance across two embodied environments. Notably, ESCA significantly reduces agent perception errors and enables open-source models to surpass proprietary baselines. We release the source code for SGCLIP model training at https://github.com/video-fm/LASER and for the embodied agent at https://github.com/video-fm/ESCA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6dcb9ca7-c4f0-49f1-bb71-377cab1213f5Cited by top-tier papers2
- RoboAgent: Chaining Basic Capabilities for Embodied Task PlanningPeiran Xu, Jiaqi Zheng, Yadong MuCVPR 2026 · 6 citations
- SegPVSG: Panoptic Video Scene Graph Generation via Temporal Focusing and Generative AugmentationYiKai Li, Quhui Ke, Jinglin Liang, Zhiyuan Zhang et al.ICML 2026
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans et al.NeurIPS 2021 · 826 citations
- ALFWorld: Aligning Text and Embodied Environments for Interactive LearningMohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk et al.ICLR 2021 · 819 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
Related papers
- LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical StudyDongil Yang, Minjin Kim, Sunghwan Kim, Beong-woo Kwak et al.ACL 2025 · 8 citations
- LASER: A Neuro-Symbolic Framework for Learning Spatio-Temporal Scene Graphs with Weak SupervisionJiani Huang, Ziyang Li, Mayur Naik, Ser-Nam LimICLR 2025
- Contrastive Localized Language-Image Pre-TrainingHong-You Chen, Zhengfeng Lai, Haotian Zhang, Xinze Wang et al.ICML 2025
- Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-Modal Structured RepresentationsYufeng Huang, Jiji Tang, Zhuo Chen, Rongsheng Zhang et al.AAAI 2024 · 65 citations
- 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene UnderstandingTatiana Zemskova, Dmitry A. YudinICCV 2025 · 7 citations
