Probing Emergent Semantics in Predictive Agents via Question Answering
Abhishek Das, Federico Carnevale, Hamza Merzic, Laura Rimell, Rosalia Schneider, Josh Abramson, Alden Hung, Arun Ahuja, Stephen Clark, Greg Wayne, Felix Hill
Abstract
Recent work has shown how predictive modeling can endow agents with rich knowledge of their surroundings, improving their ability to act in complex environments. We propose question-answering as a general paradigm to decode and understand the representations that such agents develop, applying our method to two recent approaches to predictive modeling -action-conditional CPC (Guo et al., 2018) and SimCore (Gregor et al., 2019). After training agents with these predictive objectives in a visually-rich, 3D environment with an assortment of objects, colors, shapes, and spatial configurations, we probe their internal state representations with synthetic (English) questions, without backpropagating gradients from the question-answering decoder into the agent. The performance of different agents when probed this way reveals that they learn to encode factual, and seemingly compositional, information about objects, properties and spatial relations from their physical environment. Our approach is intuitive, i.e. humans can easily interpret responses of the model as opposed to inspecting continuous vectors, and model-agnostic, i.e. applicable to any modeling approach. By revealing the implicit knowledge of objects, quantities, properties and relations acquired by agents as they learn, question-conditional agent probing can stimulate the design and development of stronger predictive learning objectives.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Auxiliary Tasks and Exploration Enable ObjectGoal NavigationJoel Ye, Dhruv Batra, Abhishek Das, Erik WijmansICCV 2021 · 137 citations
- Habitat-Web: Learning Embodied Object-Search Strategies from Human Demonstrations at ScaleRam Ramrakhya, Eric Undersander, Dhruv Batra, Abhishek DasCVPR 2022 · 73 citations
- VLM Agents Generate Their Own Memories: Distilling Experience into Embodied Programs of ThoughtGabriel Sarch, Lawrence Jang, Michael J. Tarr, William W. Cohen et al.NeurIPS 2024 · 64 citations
- Perceiving the World: Question-guided Reinforcement Learning for Text-based GamesYunqiu Xu, Meng Fang, Ling Chen, Yali Du et al.ACL 2022 · 22 citations
- EXCALIBUR: Encouraging and Evaluating Embodied ExplorationHao Zhu, Raghav Kapoor, So Yeon Min, Winson Han et al.CVPR 2023
Related papers
- Environment Predictive Coding for Visual NavigationSanthosh Kumar Ramakrishnan, Tushar Nagarajan, Ziad Al-Halah, Kristen GraumanICLR 2022 · 10 citations
- Capturing Visual Environment Structure Correlates with Control PerformanceJiahua Dong, Yunze Man, Pavel Tokmakov, Yu-Xiong WangICLR 2026 · 2 citations
- Causal-JEPA: Learning World Models through Object-Level Latent MaskingHeejeong Nam, Quentin Le Lidec, Lucas Maes, Yann LeCun et al.ICML 2026 · 7 citations
- Predict Before You Explore: Predictive Planning with Specialized Memory for Embodied Question AnsweringBowen Yuan, Sisi You, Bing-Kun BaoCVPR 2026
- Agent Modelling under Partial Observability for Deep Reinforcement LearningGeorgios Papoudakis, Filippos Christianos, Stefano V. AlbrechtNeurIPS 2021 · 110 citations
