Quantum Policy Gradient Algorithm with Optimized Action Decoding
Nico Meyer, Daniel D. Scherer, Axel Plinge, Christopher Mutschler, Michael J. Hartmann
Abstract
Quantum machine learning implemented by variational quantum circuits (VQCs) is considered a promising concept for the noisy intermediate-scale quantum computing era. Focusing on applications in quantum reinforcement learning, we propose a specific action decoding procedure for a quantum policy gradient approach. We introduce a novel quality measure that enables us to optimize the classical post-processing required for action selection, inspired by local and global quantum measurements. The resulting algorithm demonstrates a significant performance improvement in several benchmark environments. With this technique, we successfully execute a full training routine on a 5-qubit hardware device. Our method introduces only negligible classical overhead and has the potential to improve VQC-based algorithms beyond the field of quantum reinforcement learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext abacd59a-da2b-44f8-8127-4d321f28b799Cited by top-tier papers2
- Predictive Performance of Deep Quantum Data Re-uploading ModelsXin Wang, Hanxiao Tao, Rebing WuICML 2025
- Benchmarking Quantum Reinforcement LearningNico Meyer, Christian Ufrecht, George Yammine, Georgios D. Kontes et al.ICML 2025
Builds on1
Related papers
- Curriculum reinforcement learning for quantum architecture search under hardware errorsYash J. Patel, Akash Kundu, Mateusz Ostaszewski, Xavier Bonet-Monroig et al.ICLR 2024 · 54 citations
- Offline Quantum Reinforcement Learning in a Conservative MannerZhihao Cheng, Kaining Zhang, Li Shen, Dacheng TaoAAAI 2023 · 7 citations
- TensorRL-QAS: Reinforcement learning with tensor networks for improved quantum architecture searchAkash Kundu, Stefano ManginiNeurIPS 2025 · 9 citations
- Learning to Optimize Variational Quantum Circuits to Solve Combinatorial ProblemsSami Khairy, Ruslan Shaydulin, Lukasz Cincio, Yuri Alexeev et al.AAAI 2020 · 155 citations
- Reinforcement learning for optimization of variational quantum circuit architecturesMateusz Ostaszewski, Lea M. Trenkwalder, Wojciech Masarczyk, Eleanor Scerri et al.NeurIPS 2021 · 204 citations
