Policy Learning Using Weak Supervision
Jingkang Wang, Hongyi Guo, Zhaowei Zhu, Yang Liu
Abstract
Most existing policy learning solutions require the learning agents to receive highquality supervision signals such as well-designed rewards in reinforcement learning (RL) or high-quality expert demonstrations in behavioral cloning (BC). These quality supervisions are usually infeasible or prohibitively expensive to obtain in practice. We aim for a unified framework that leverages the available cheap weak supervisions to perform policy learning efficiently. To handle this problem, we treat the "weak supervision" as imperfect information coming from a peer agent, and evaluate the learning agent's policy based on a "correlated agreement" with the peer agent's policy (instead of simple agreements). Our approach explicitly punishes a policy for overfitting to the weak supervision. In addition to theoretical guarantees, extensive evaluations on tasks including RL with noisy rewards, BC with weak demonstrations, and standard policy co-training show that our method leads to substantial performance improvements, especially when the complexity or the noise of the learning environments is high. Weak Demonstrations
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e1d6abb-6772-490d-bf55-375f2aef9f91Cited by top-tier papers7
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu et al.ICLR 2022 · 338 citations
- Detecting Corrupted Labels Without Training a Model to PredictZhaowei Zhu, Zihao Dong, Yang LiuICML 2022 · 84 citations
- Open-Sampling: Exploring Out-of-Distribution data for Re-balancing Long-tailed datasetsHongxin Wei, Lue Tao, Renchunzi Xie, Lei Feng et al.ICML 2022 · 46 citations
- Beyond Images: Label Noise Transition Matrix Estimation for Tasks with Lower-Quality FeaturesZhaowei Zhu, Jialu Wang, Yang LiuICML 2022 · 43 citations
- To Aggregate or Not? Learning with Separate Noisy LabelsJiaheng Wei, Zhaowei Zhu, Tianyi Luo, Ehsan Amid et al.KDD 2023 · 21 citations
Builds on10
- Peer Loss Functions: Learning from Noisy Labels without Knowing Noise RatesYang Liu, Hongyi GuoICML 2020 · 280 citations
- Provably End-to-end Label-noise Learning without Anchor PointsXuefeng Li, Tongliang Liu, Bo Han, Gang Niu et al.ICML 2021 · 161 citations
- Reinforcement Learning with Perturbed RewardsJingkang Wang, Yang Liu, Bo LiAAAI 2020 · 161 citations
- Clusterability as an Alternative to Anchor Points When Learning with Noisy LabelsZhaowei Zhu, Yiwen Song, Yang LiuICML 2021 · 112 citations
- Behavioral Cloning from Noisy DemonstrationsFumihiro Sasaki, Ryota YamashinaICLR 2021 · 94 citations
Related papers
- Guided Policy Optimization under Partial ObservabilityYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 4 citations
- SQIL: Imitation Learning via Reinforcement Learning with Sparse RewardsSiddharth Reddy, Anca D. Dragan, Sergey LevineICLR 2020 · 299 citations
- Should I Run Offline Reinforcement Learning or Behavioral Cloning?Aviral Kumar, Joey Hong, Anikait Singh, Sergey LevineICLR 2022 · 84 citations
- Co-learning: Learning from Noisy Labels with Self-supervisionCheng Tan, Jun Xia, Lirong Wu, Stan Z. LiACM MM 2021 · 145 citations
- Learning from Interventions Using Hierarchical Policies for Safe LearningJing Bi, Vikas Dhiman, Tianyou Xiao, Chenliang XuAAAI 2020 · 9 citations
