RLang: A Declarative Language for Describing Partial World Knowledge to Reinforcement Learning Agents
Rafael Rodríguez-Sánchez, Benjamin Adin Spiegel, Jennifer Wang, Roma Patel, Stefanie Tellex, George Konidaris
Abstract
We introduce RLang, a domain-specific language (DSL) for communicating domain knowledge to an RL agent. Unlike existing RL DSLs that ground to single elements of a decision-making formalism (e.g., the reward function or policy), RLang can specify information about every element of a Markov decision process. We define precise syntax and grounding semantics for RLang, and provide a parser that grounds RLang programs to an algorithm-agnostic partial world model and policy that can be exploited by an RL agent. We provide a series of example RLang programs demonstrating how different RL methods can exploit the resulting knowledge, encompassing model-free and model-based tabular algorithms, policy gradient and value-based methods, hierarchical approaches, and deep methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58cd12e6-2013-4a91-a659-78d1c3586613Cited by top-tier papers3
- WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the EnvironmentHao Tang, Darren Key, Kevin EllisNeurIPS 2024 · 123 citations
- InstructFlow: Adaptive Symbolic Constraint-Guided Code Generation for Long-Horizon PlanningHaotian Chi, Zeyu Feng, Yueming Lyu, Chengqi Zheng et al.NeurIPS 2025 · 6 citations
- Reward Learning from Multiple Feedback TypesYannick Metz, András Geiszl, Raphaël Baur, Mennatallah El-AssadyICLR 2025
Builds on3
- Learning Rewards From Linguistic FeedbackTheodore R. Sumers, Mark K. Ho, Robert X. D. Hawkins, Karthik Narasimhan et al.AAAI 2021 · 67 citations
- Program Guided AgentShao-Hua Sun, Te-Lin Wu, Joseph J. LimICLR 2020 · 63 citations
- RTFM: Generalising to New Environment Dynamics via ReadingVictor Zhong, Tim Rocktäschel, Edward GrefenstetteICLR 2020 · 44 citations
Related papers
- Reward Machines for Deep RL in Noisy and Uncertain EnvironmentsAndrew C. Li, Zizhao Chen, Toryn Q. Klassen, Pashootan Vaezipoor et al.NeurIPS 2024 · 19 citations
- LTL2Action: Generalizing LTL Instructions for Multi-Task RLPashootan Vaezipoor, Andrew C. Li, Rodrigo Toro Icarte, Sheila A. McIlraithICML 2021 · 106 citations
- Tell me why! Explanations support learning relational and causal structureAndrew K. Lampinen, Nicholas A. Roy, Ishita Dasgupta, Stephanie C. Y. Chan et al.ICML 2022 · 51 citations
- Learning to Model the World With LanguageJessy Lin, Yuqing Du, Olivia Watkins, Danijar Hafner et al.ICML 2024 · 76 citations
- Improving Policy Learning via Language Dynamics DistillationVictor Zhong, Jesse Mu, Luke Zettlemoyer, Edward Grefenstette et al.NeurIPS 2022 · 16 citations
