Policy Learning Using Weak Supervision
Jingkang Wang, Hongyi Guo, Zhaowei Zhu, Yang Liu
摘要
Most existing policy learning solutions require the learning agents to receive highquality supervision signals such as well-designed rewards in reinforcement learning (RL) or high-quality expert demonstrations in behavioral cloning (BC). These quality supervisions are usually infeasible or prohibitively expensive to obtain in practice. We aim for a unified framework that leverages the available cheap weak supervisions to perform policy learning efficiently. To handle this problem, we treat the "weak supervision" as imperfect information coming from a peer agent, and evaluate the learning agent's policy based on a "correlated agreement" with the peer agent's policy (instead of simple agreements). Our approach explicitly punishes a policy for overfitting to the weak supervision. In addition to theoretical guarantees, extensive evaluations on tasks including RL with noisy rewards, BC with weak demonstrations, and standard policy co-training show that our method leads to substantial performance improvements, especially when the complexity or the noise of the learning environments is high. Weak Demonstrations
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu 等ICLR 2022 · 被引用 338 次
- Detecting Corrupted Labels Without Training a Model to PredictZhaowei Zhu, Zihao Dong, Yang LiuICML 2022 · 被引用 84 次
- Open-Sampling: Exploring Out-of-Distribution data for Re-balancing Long-tailed datasetsHongxin Wei, Lue Tao, Renchunzi Xie, Lei Feng 等ICML 2022 · 被引用 46 次
- Beyond Images: Label Noise Transition Matrix Estimation for Tasks with Lower-Quality FeaturesZhaowei Zhu, Jialu Wang, Yang LiuICML 2022 · 被引用 43 次
- To Aggregate or Not? Learning with Separate Noisy LabelsJiaheng Wei, Zhaowei Zhu, Tianyi Luo, Ehsan Amid 等KDD 2023 · 被引用 21 次
它引用的顶会 Paper10
- Peer Loss Functions: Learning from Noisy Labels without Knowing Noise RatesYang Liu, Hongyi GuoICML 2020 · 被引用 280 次
- Provably End-to-end Label-noise Learning without Anchor PointsXuefeng Li, Tongliang Liu, Bo Han, Gang Niu 等ICML 2021 · 被引用 161 次
- Reinforcement Learning with Perturbed RewardsJingkang Wang, Yang Liu, Bo LiAAAI 2020 · 被引用 161 次
- Clusterability as an Alternative to Anchor Points When Learning with Noisy LabelsZhaowei Zhu, Yiwen Song, Yang LiuICML 2021 · 被引用 112 次
- Behavioral Cloning from Noisy DemonstrationsFumihiro Sasaki, Ryota YamashinaICLR 2021 · 被引用 94 次
相关 Paper
- Guided Policy Optimization under Partial ObservabilityYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 被引用 4 次
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 被引用 299 次
- Should I Run Offline Reinforcement Learning or Behavioral Cloning?Aviral Kumar, Joey Hong, Anikait Singh, Sergey LevineICLR 2022 · 被引用 84 次
- Co-learning: Learning from Noisy Labels with Self-supervisionCheng Tan, Jun Xia, Lirong Wu, Stan Z. LiACM MM 2021 · 被引用 145 次
- Learning from Interventions Using Hierarchical Policies for Safe LearningJing Bi, Vikas Dhiman, Tianyou Xiao, Chenliang XuAAAI 2020 · 被引用 9 次
