Beyond Black-Box Advice: Learning-Augmented Algorithms for MDPs with Q-Value Predictions
Tongxin Li, Yiheng Lin, Shaolei Ren, Adam Wierman
Abstract
We study the tradeoff between consistency and robustness in the context of a single-trajectory time-varying Markov Decision Process (MDP) with untrusted machine-learned advice. Our work departs from the typical approach of treating advice as coming from black-box sources by instead considering a setting where additional information about how the advice is generated is available. We prove a first-of-its-kind consistency and robustness tradeoff given Q-value advice under a general MDP model that includes both continuous and discrete state/action spaces. Our results highlight that utilizing Q-value advice enables dynamic pursuit of the better of machine-learned advice and a robust baseline, thus result in near-optimal performance guarantees, which provably improves what can be obtained solely with black-box advice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96bcb1e5-5c29-4a14-a555-67abd122d3d2Cited by top-tier papers4
- Safe Exploitative Play with Untrusted Type BeliefsTongxin Li, Tinashe Handina, Shaolei Ren, Adam WiermanNeurIPS 2024 · 3 citations
- Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen ApproachChenbei Lu, Zaiwei Chen, Tongxin Li, Chenye Wu et al.NeurIPS 2025 · 3 citations
- Disentangling Linear Quadratic Control with Untrusted ML PredictionsTongxin Li, Hao Liu, Yisong YueNeurIPS 2024 · 3 citations
- Efficiently Solving Discounted MDPs via Predictions with Unknown Prediction ErrorsLixing Lyu, Jiashuo Jiang, Wang Chi CheungICML 2026
Builds on16
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- The Primal-Dual method for Learning Augmented AlgorithmsÉtienne Bamas, Andreas Maggiori, Ola SvenssonNeurIPS 2020 · 171 citations
- Online metric algorithms with untrusted predictionsAntonios Antoniadis, Christian Coester, Marek Eliás, Adam Polak et al.ICML 2020 · 170 citations
- Optimal Robustness-Consistency Trade-offs for Learning-Augmented Online AlgorithmsAlexander Wei, Fred ZhangNeurIPS 2020 · 129 citations
- Near-Optimal Bounds for Online Caching with Machine Learned AdviceDhruv RohatgiSODA 2020 · 88 citations
Related papers
- Advice-Guided Reinforcement Learning in a non-Markovian EnvironmentDaniel Neider, Jean-Raphaël Gaglione, Ivan Gavran, Ufuk Topcu et al.AAAI 2021 · 40 citations
- Online Robust Reinforcement Learning with Model UncertaintyYue Wang, Shaofeng ZouNeurIPS 2021 · 157 citations
- Efficient Solution and Learning of Robust Factored MDPsYannik Schnitzer, Alessandro Abate, David ParkerAAAI 2026 · 1 citation
- Single-Trajectory Distributionally Robust Reinforcement LearningZhipeng Liang, Xiaoteng Ma, José H. Blanchet, Jun Yang et al.ICML 2024 · 15 citations
- Time-Constrained Robust MDPsAdil Zouitine, David Bertoin, Pierre Clavier, Matthieu Geist et al.NeurIPS 2024 · 6 citations
