Diagnosis, Feedback, Adaptation: A Human-in-the-Loop Framework for Test-Time Policy Adaptation
Andi Peng, Aviv Netanyahu, Mark K. Ho, Tianmin Shu, Andreea Bobu, Julie Shah, Pulkit Agrawal
摘要
Policies often fail due to distribution shift -- changes in the state and reward that occur when a policy is deployed in new environments. Data augmentation can increase robustness by making the model invariant to task-irrelevant changes in the agent's observation. However, designers don't know which concepts are irrelevant a priori, especially when different end users have different preferences about how the task is performed. We propose an interactive framework to leverage feedback directly from the user to identify personalized task-irrelevant concepts. Our key idea is to generate counterfactual demonstrations that allow users to quickly identify possible task-relevant and irrelevant concepts. The knowledge of task-irrelevant concepts is then used to perform data augmentation and thus obtain a policy adapted to personalized user objectives. We present experiments validating our framework on discrete and continuous control tasks with real human users. Our method (1) enables users to better understand agent failure, (2) reduces the number of demonstrations required for fine-tuning, and (3) aligns the agent to individual user task preferences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning with Language-Guided State AbstractionsAndi Peng, Ilia Sucholutsky, Belinda Z. Li, Theodore R. Sumers 等ICLR 2024 · 被引用 20 次
- PRIMT: Preference-based Reinforcement Learning with Multimodal Feedback and Trajectory Synthesis from Foundation ModelsRuiqi Wang, Dezhong Zhao, Ziqin Yuan, Tianyu Shao 等NeurIPS 2025 · 被引用 9 次
- Pragmatic Feature Preferences: Learning Reward-Relevant Preferences from Human InputAndi Peng, Yuying Sun, Tianmin Shu, David AbelICML 2024 · 被引用 7 次
- Test-Time Adaptation for Online Vision-Language Navigation with Feedback-based Reinforcement LearningSungjune Kim, Gyeongrok Oh, Heeju Ko, Daehyun Ji 等ICML 2025
它引用的顶会 Paper2
相关 Paper
- Causal Action Influence Aware Counterfactual Data AugmentationNúria Armengol Urpí, Marco Bagatella, Marin Vlastelica, Georg MartiusICML 2024 · 被引用 11 次
- Active Fine-Tuning of Multi-Task PoliciesMarco Bagatella, Jonas Hübotter, Georg Martius, Andreas KrauseICML 2025
- Automatic Data Augmentation for Generalization in Reinforcement LearningRoberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov 等NeurIPS 2021 · 被引用 143 次
- C2L: Causally Contrastive Learning for Robust Text ClassificationSeungtaek Choi, Myeongho Jeong, Hojae Han, Seung-won HwangAAAI 2022 · 被引用 52 次
- Making Learner Weakness Actionable for Learning from Demonstration with Novice TeachersYuqing Zhu, Matthew HowardICML 2026
