Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning
Tongxi Wang, Zhuoyang Xia, Xinran Chen, Shan Liu
Abstract
Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entropy, causing over-exploration during stable periods and under-exploration after drift (thus slow recovery), and leaving unanswered the principled question of how exploration intensity should scale with drift magnitude. We show that, under standard assumptions, entropy scheduling in non-stationary maximum-entropy RL can be cast as the dynamic-regret trade-off between tracking a drifting comparator and stabilizing updates, yielding a square-root scaling rule for the entropy weight in terms of a (possibly conservative) online non-stationarity proxy. Building on this, we propose AES (Adaptive Entropy Scheduling), which adaptively adjusts the entropy coefficient/temperature online using observable drift proxies during training, requiring almost no structural changes and incurring minimal overhead. Across 4 algorithm variants, 12 tasks, and 4 drift modes, AES significantly reduces the fraction of performance degradation caused by drift and accelerates recovery after abrupt changes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 00b03412-99d9-418d-9edf-e01b9c6dd771Cited by top-tier papers3
- ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Mingyu Zhang et al.CVPR 2026 · 16 citations
- Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image RetrievalZhiheng Fu, Yupeng Hu, Qianyun Yang, Shiqi Zhang et al.CVPR 2026 · 16 citations
- TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image RetrievalZixu Li, Yupeng Hu, Zhiheng Fu, Zhiwei Chen et al.ACL 2026 · 13 citations
Builds on7
- Deep Reinforcement Learning amidst Continual Structured Non-StationarityAnnie Xie, James Harrison, Chelsea FinnICML 2021 · 43 citations
- VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector AnimationGuotao Liang, Zhangcheng Wang, Chuang Wang, Juncheng Hu et al.ICML 2026 · 11 citations
- Pathology-Aware Prototype Evolution via LLM-Driven Semantic Disambiguation for Multicenter Diabetic Retinopathy DiagnosisChunzheng Zhu, Yangfang Lin, Jialin Shao, Jianxin Lin et al.ACM MM 2025 · 6 citations
- Multi-Object Sketch Animation with Grouping and Motion Trajectory PriorsGuotao Liang, Juncheng Hu, Ximing Xing, Jing Zhang et al.ACM MM 2025 · 4 citations
- Convergence Theorems for Entropy-Regularized and Distributional Reinforcement LearningYash Jhaveri, Harley Wiltzer, Patrick Shafto, Marc G. Bellemare et al.NeurIPS 2025 · 3 citations
Related papers
- Demystifying Entropy Control in LLM RL Training: Theoretical Analysis and Dynamic SchedulingJingchu Gai, Guanning Zeng, Huaqing ZHANG, Han Zhong et al.ICML 2026
- An Adaptive Entropy-Regularization Framework for Multi-Agent Reinforcement LearningWoojun Kim, Youngchul SungICML 2023 · 20 citations
- Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RLGuojian Zhan, Likun Wang, Pengcheng Wang, Feihong Zhang et al.ICML 2026
- When Maximum Entropy Misleads Policy OptimizationRuipeng Zhang, Ya-Chien Chang, Sicun GaoICML 2025
- Stable Deep Reinforcement Learning via Isotropic Gaussian RepresentationsAli Saheb pasand, Johan Obando-Ceron, Aaron Courville, Pouya Bashivan et al.ICML 2026 · 5 citations
