Belief-State Query Policies for User-Aligned POMDPs
Daniel Bramblett, Siddharth Srivastava
摘要
Planning in real-world settings often entails addressing partial observability while aligning with users' requirements. We present a novel framework for expressing users' constraints and preferences about agent behavior in a partially observable setting using parameterized belief-state query (BSQ) policies in the setting of goal-oriented partially observable Markov decision processes (gPOMDPs). We present the first formal analysis of such constraints and prove that while the expected cost function of a parameterized BSQ policy w.r.t its parameters is not convex, it is piecewise constant and yields an implicit discrete parameter search space that is finite for finite horizons. This theoretical result leads to novel algorithms that optimize gPOMDP agent behavior with guaranteed user alignment. Analysis proves that our algorithms converge to the optimal user-aligned behavior in the limit. Empirical results show that parameterized BSQ policies provide a computationally feasible approach for user-aligned planning in partially observable settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- LTL2Action: Generalizing LTL Instructions for Multi-Task RLPashootan Vaezipoor, Andrew C. Li, Rodrigo Toro Icarte, Sheila A. McIlraithICML 2021 · 被引用 106 次
- The Perils of Trial-and-Error Reward Design: Misdesign through Overfitting and Invalid Task SpecificationsSerena Booth, W. Bradley Knox, Julie Shah, Scott Niekum 等AAAI 2023 · 被引用 103 次
- Explicable Reward Design for Reinforcement Learning AgentsRati Devidze, Goran Radanovic, Parameswaran Kamalaruban, Adish SinglaNeurIPS 2021 · 被引用 60 次
- Policy Optimization with Linear Temporal Logic ConstraintsCameron Voloshin, Hoang Minh Le, Swarat Chaudhuri, Yisong YueNeurIPS 2022 · 被引用 28 次
- Behavior Alignment via Reward Function OptimizationDhawal Gupta, Yash Chandak, Scott M. Jordan, Philip S. Thomas 等NeurIPS 2023 · 被引用 27 次
相关 Paper
- Inferring Implicit Goals Across Differing Task ModelsSilvia Tulli, Stylianos Loukas Vasileiou, Mohamed Chetouani, Sarath SreedharanAAAI 2026
- AI Alignment with Changing and Influenceable Reward FunctionsMicah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell 等ICML 2024 · 被引用 44 次
- The Value Function Semi-Algebraic Set in Partially Observable Markov Decision ProcessesRyan Anderson, Guido MontufarICML 2026
- Robust Finite-State Controllers for Uncertain POMDPsMurat Cubuktepe, Nils Jansen, Sebastian Junges, Ahmadreza Marandi 等AAAI 2021 · 被引用 35 次
- What Should Be Observed for Optimal Reward in POMDPs?Alyzia-Maria Konsta, Alberto Lluch-Lafuente, Christoph MathejaCAV 2024 · 被引用 2 次
