q-exponential family for policy optimization
Lingwei Zhu, Haseeb Shah, Han Wang, Yukie Nagai, Martha White
摘要
Policy optimization methods benefit from a simple and tractable policy parametrization, usually the Gaussian for continuous action spaces. In this paper, we consider a broader policy family that remains tractable: the q-exponential family. This family of policies is flexible, allowing the specification of both heavy-tailed policies (q > 1) and light-tailed policies (q < 1). This paper examines the interplay between q-exponential policies for several actor-critic algorithms conducted on both online and offline problems. We find that heavy-tailed policies are more effective in general and can consistently improve on Gaussian. In particular, we find the Student's t-distribution to be more stable than the Gaussian across settings and that a heavy-tailed q-Gaussian for Tsallis Advantage Weighted Actor-Critic consistently performs well in offline benchmark problems. In summary, we find that the Student's t policy a strong candidate for drop-in replacement to the Gaussian. Our code is available at https://github.com/lingweizhu/qexp .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- A Policy-Guided Imitation Approach for Offline Reinforcement LearningHaoran Xu, Li Jiang, Jianxiong Li, Xianyuan ZhanNeurIPS 2022 · 被引用 86 次
- Offline RL with No OOD Actions: In-Sample Learning via Implicit Value RegularizationHaoran Xu, Li Jiang, Jianxiong Li, Zhuoran Yang 等ICLR 2023 · 被引用 4 次
相关 Paper
- Fat-to-Thin Policy Optimization: Offline Reinforcement Learning with Sparse PoliciesLingwei Zhu, Han Wang, Yukie NagaiICLR 2025
- On the Hidden Biases of Policy Mirror Ascent in Continuous Action SpacesAmrit Singh Bedi, Souradip Chakraborty, Anjaly Parayil, Brian M. Sadler 等ICML 2022 · 被引用 20 次
- Tackling Heavy-Tailed Q-Value Bias in Offline-to-Online Reinforcement Learning with Laplace-Robust ModelingRuibo Guo, Lei Liu, Rui Yang, Junjie Shen 等ICLR 2026
- Provably Robust Temporal Difference Learning for Heavy-Tailed RewardsSemih Cayci, Atilla EryilmazNeurIPS 2023 · 被引用 12 次
- General Munchausen Reinforcement Learning with Tsallis Kullback-Leibler DivergenceLingwei Zhu, Zheng Chen, Matthew Schlegel, Martha WhiteNeurIPS 2023 · 被引用 4 次
