Asking the Right Questions: Learning Interpretable Action Models Through Query Answering
Pulkit Verma, Shashank Rao Marpally, Siddharth Srivastava
Abstract
This paper develops a new approach for estimating an interpretable, relational model of a black-box autonomous agent that can plan and act. Our main contributions are a new paradigm for estimating such models using a rudimentary query interface with the agent and a hierarchical querying algorithm that generates an interrogation policy for estimating the agent's internal model in a user-interpretable vocabulary. Empirical evaluation of our approach shows that despite the intractable search space of possible agent models, our approach allows correct and scalable estimation of interpretable agent models for a wide class of black-box autonomous agents. Our results also show that this approach can use predicate classifiers to learn interpretable models of planning agents that represent states as images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Differential Assessment of Black-Box AI AgentsRashmeet Kaur Nayyar, Pulkit Verma, Siddharth SrivastavaAAAI 2022 · 19 citations
- Autonomous Capability Assessment of Sequential Decision-Making Systems in Stochastic SettingsPulkit Verma, Rushang Karia, Siddharth SrivastavaNeurIPS 2023 · 14 citations
Builds on1
Related papers
- Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model ExplanationThien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun SakumaAAAI 2022 · 11 citations
- TripleTree: A Versatile Interpretable Representation of Black Box Agents and their EnvironmentsTom Bewley, Jonathan LawryAAAI 2021 · 33 citations
- ProtoX: Explaining a Reinforcement Learning Agent via PrototypingRonilo J. Ragodos, Tong Wang, Qihang Lin, Xun ZhouNeurIPS 2022 · 13 citations
- Generating High-Quality Explanations for Navigation in Partially-Revealed EnvironmentsGregory J. SteinNeurIPS 2021 · 19 citations
- Probing Emergent Semantics in Predictive Agents via Question AnsweringAbhishek Das, Federico Carnevale, Hamza Merzic, Laura Rimell et al.ICML 2020 · 18 citations
