Stay Moral and Explore: Learn to Behave Morally in Text-based Games
Zijing Shi, Meng Fang, Yunqiu Xu, Ling Chen, Yali Du
Abstract
Reinforcement learning (RL) in text-based games has developed rapidly and achieved promising results. However, little effort has been expended to design agents that pursue objectives while behaving morally, which is a critical issue in the field of autonomous agents. In this paper, we propose a general framework named Moral Awareness Adaptive Learning (MorAL) that enhances the morality capacity of an agent using a plugin moral-aware learning model. The framework allows the agent to execute task learning and morality learning adaptively. The agent selects trajectories from past experiences during task learning. Meanwhile, the trajectories are used to conduct self-imitation learning with a moral-enhanced objective. In order to achieve the trade-off between morality and task progress, the agent uses the combination of task policy and moral policy for action selection. We evaluate on the Jiminy Cricket benchmark, a set of text-based games with various scenes and dense morality annotations. Our experiments demonstrate that, compared with strong contemporary value alignment approaches, the proposed framework improves task performance while reducing immoral behaviours in various games.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6543305c-99a7-4f0d-ae40-790e576f377cCited by top-tier papers5
- Large Language Models Are Neurosymbolic ReasonersMeng Fang, Shilong Deng, Yudi Zhang, Zijing Shi et al.AAAI 2024 · 53 citations
- CHBias: Bias Evaluation and Mitigation of Chinese Conversational Language ModelsJiaxu Zhao, Meng Fang, Zijing Shi, Yitong Li et al.ACL 2023 · 11 citations
- Monte Carlo Planning with Large Language Model for Text-Based Game AgentsZijing Shi, Meng Fang, Ling ChenICLR 2025
- Understanding Large Language Model Vulnerabilities to Social Bias AttacksJiaxu Zhao, Meng Fang, Fanghua Ye, Ke Xu et al.ACL 2025
- Unmasking Style Sensitivity: A Causal Analysis of Bias Evaluation Instability in Large Language ModelsJiaxu Zhao, Meng Fang, Kun Zhang, Mykola PechenizkiyACL 2025
Related papers
- Learning Human-like Representations to Enable Learning Human ValuesAndrea Wynn, Ilia Sucholutsky, Tom GriffithsNeurIPS 2024 · 11 citations
- Learning to Follow Instructions in Text-Based GamesMathieu Tuli, Andrew C. Li, Pashootan Vaezipoor, Toryn Q. Klassen et al.NeurIPS 2022 · 21 citations
- Act-Adaptive Margin: Dynamically Calibrating Reward Models for Subjective AmbiguityFeiteng Fang, Dingwei Chen, Xiang Huang, Ting-En Lin et al.ACL 2026 · 3 citations
- Deep Reinforcement Learning with Stacked Hierarchical Attention for Text-based GamesYunqiu Xu, Meng Fang, Ling Chen, Yali Du et al.NeurIPS 2020 · 48 citations
- ROMA: Multi-Agent Reinforcement Learning with Emergent RolesTonghan Wang, Heng Dong, Victor R. Lesser, Chongjie ZhangICML 2020 · 286 citations
