TW-CRL: Time-Weighted Contrastive Reward Learning for Efficient Inverse Reinforcement Learning
Yuxuan Li, Yicheng Gao, Ning Yang, Stephen Xia
Abstract
Episodic tasks in Reinforcement Learning (RL) often pose challenges due to sparse reward signals and high-dimensional state spaces, which hinder efficient learning. Additionally, these tasks often feature hidden “trap states”—irreversible failures that prevent task completion but do not provide explicit negative rewards to guide agents away from repeated errors. To address these issues, we propose Time-Weighted Contrastive Reward Learning (TW-CRL), an Inverse Reinforcement Learning (IRL) framework that leverages both successful and failed demonstrations. By incorporating temporal information, TW-CRL learns a dense reward function that identifies critical states associated with success or failure. This approach not only enables agents to avoid trap states but also encourages meaningful exploration beyond simple imitation of expert trajectories. Empirical evaluations on navigation tasks and robotic manipulation benchmarks demonstrate that TW-CRL surpasses state-of-the-art methods, achieving improved efficiency and robustness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- A Unifying View of Optimism in Episodic Reinforcement LearningGergely Neu, Ciara Pike-BurkeNeurIPS 2020 · 79 citations
- Dense Reward for Free in Reinforcement Learning from Human FeedbackAlex James Chan, Hao Sun, Samuel Holt, Mihaela van der SchaarICML 2024 · 74 citations
- Risk-Sensitive Reinforcement Learning with Function Approximation: A Debiasing ApproachYingjie Fei, Zhuoran Yang, Zhaoran WangICML 2021 · 53 citations
- How to Stay Curious while avoiding Noisy TVs using Aleatoric Uncertainty EstimationAugustine N. Mavor-Parker, Kimberly A. Young, Caswell Barry, Lewis D. GriffinICML 2022 · 32 citations
- RA-PbRL: Provably Efficient Risk-Aware Preference-Based Reinforcement LearningYujie Zhao, Jose E. Aguilar Escamilla, Weyl Lu, Huazheng WangNeurIPS 2024 · 11 citations
Related papers
- Multi-Modal Inverse Constrained Reinforcement Learning from a Mixture of DemonstrationsGuanren Qiao, Guiliang Liu, Pascal Poupart, Zhiqiang XuNeurIPS 2023 · 28 citations
- Learning Shared Safety Constraints from Multi-task DemonstrationsKonwoo Kim, Gokul Swamy, Zuxin Liu, Ding Zhao et al.NeurIPS 2023 · 31 citations
- BC-IRL: Learning Generalizable Reward Functions from DemonstrationsAndrew Szot, Amy Zhang, Dhruv Batra, Zsolt Kira et al.ICLR 2023 · 1 citation
- Provably Efficient Exploration in Inverse Constrained Reinforcement LearningBo Yue, Jian Li, Guiliang LiuICML 2025
- Understanding Constraint Inference in Safety-Critical Inverse Reinforcement LearningBo Yue, Shufan Wang, Ashish Gaurav, Jian Li et al.ICLR 2025
