Safe Reinforcement Learning with Natural Language Constraints
Tsung-Yen Yang, Michael Y. Hu, Yinlam Chow, Peter J. Ramadge, Karthik Narasimhan
摘要
In this paper, we tackle the problem of learning control policies for tasks when provided with constraints in natural language. In contrast to instruction following, language here is used not to specify goals, but rather to describe situations that an agent must avoid during its exploration of the environment. Specifying constraints in natural language also differs from the predominant paradigm in safe reinforcement learning, where safety criteria are enforced by hand-defined cost functions. While natural language allows for easy and flexible specification of safety constraints and budget limitations, its ambiguous nature presents a challenge when mapping these specifications into representations that can be used by techniques for safe reinforcement learning. To address this, we develop a model that contains two components: (1) a constraint interpreter to encode natural language constraints into vector representations capturing spatial and temporal information on forbidden states, and (2) a policy network that uses these representations to output a policy with minimal constraint violations. Our model is end-to-end differentiable and we train it using a recently proposed algorithm for constrained policy optimization. To empirically demonstrate the effectiveness of our approach, we create a new benchmark task for autonomous navigation with crowd-sourced free-form text specifying three different types of constraints. Our method outperforms several baselines by achieving 6-7 times higher returns and 76% fewer constraint violations on average. Dataset and code to reproduce our experiments are available at this https URL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- GLUECons: A Generic Benchmark for Learning under ConstraintsHossein Rajaby Faghihi, Aliakbar Nafar, Chen Zheng, Roshanak Mirzaee 等AAAI 2023 · 被引用 18 次
- Embedding-Aligned Language ModelsGuy Tennenholtz, Yinlam Chow, Chih-Wei Hsu, Lior Shani 等NeurIPS 2024 · 被引用 7 次
- From Text to Trajectory: Exploring Complex Constraint Representation and Decomposition in Safe Reinforcement LearningPusen Dong, Tianchen Zhu, Yue Qiu, Haoyi Zhou 等NeurIPS 2024 · 被引用 2 次
它引用的顶会 Paper5
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 被引用 403 次
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 被引用 306 次
- Safe Reinforcement Learning via Curriculum InductionMatteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause 等NeurIPS 2020 · 被引用 109 次
- Learning to Follow Directions in Street ViewKarl Moritz Hermann, Mateusz Malinowski, Piotr Mirowski, Andras Banki-Horvath 等AAAI 2020 · 被引用 78 次
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingWeituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin 等CVPR 2020
相关 Paper
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 被引用 58 次
- Safety Representations for Safer Policy LearningKaustubh Mani, Vincent Mai, Charlie Gauthier, Annie S. Chen 等ICLR 2025
- First Order Constrained Optimization in Policy SpaceYiming Zhang, Quan Vuong, Keith W. RossNeurIPS 2020 · 被引用 238 次
- Accelerating Safe Reinforcement Learning with Constraint-mismatched Baseline PoliciesTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICML 2021 · 被引用 20 次
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 被引用 127 次
