Entropy Regularization for Population Estimation
Ben Chugg, Peter Henderson, Jacob S. Goldin, Daniel E. Ho
Abstract
Entropy regularization is known to improve exploration in sequential decision-making problems. We show that this same mechanism can also lead to nearly unbiased and lower-variance estimates of the mean reward in the optimize-and-estimate structured bandit setting. Mean reward estimation (i.e., population estimation) tasks have recently been shown to be essential for public policy settings where legal constraints often require precise estimates of population metrics. We show that leveraging entropy and KL divergence can yield a better trade-off between reward and estimator variance than existing baselines, all while remaining nearly unbiased. These properties of entropy regularization illustrate an exciting potential for bringing together the optimal exploration and estimation literature.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Inference for Batched BanditsKelly W. Zhang, Lucas Janson, Susan A. MurphyNeurIPS 2020 · 115 citations
- Post-Contextual-Bandit InferenceAurélien Bibaut, Maria Dimakopoulou, Nathan Kallus, Antoine Chambaz et al.NeurIPS 2021 · 58 citations
- Online Multi-Armed Bandits with Adaptive InferenceMaria Dimakopoulou, Zhimei Ren, Zhengyuan ZhouNeurIPS 2021 · 47 citations
- On Conditional Versus Marginal Bias in Multi-Armed BanditsJaehyeok Shin, Aaditya Ramdas, Alessandro RinaldoICML 2020 · 13 citations
- Integrating Reward Maximization and Population Estimation: Sequential Decision-Making for Internal Revenue Service Audit SelectionPeter Henderson, Ben Chugg, Brandon R. Anderson, Kristen M. Altenburger et al.AAAI 2023 · 10 citations
Related papers
- Intrinsic Benefits of Categorical Distributional Loss: Uncertainty-aware Regularized Exploration in Reinforcement LearningKe Sun, Yingnan Zhao, Enze Shi, Yafei Wang et al.NeurIPS 2025 · 1 citation
- Promoting Stochasticity for Expressive Policies via a Simple and Efficient Regularization MethodQi Zhou, Yufei Kuang, Zherui Qiu, Houqiang Li et al.NeurIPS 2020 · 9 citations
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin et al.NeurIPS 2020 · 106 citations
- Iterative Amortized Policy OptimizationJoseph Marino, Alexandre Piché, Alessandro Davide Ialongo, Yisong YueNeurIPS 2021 · 27 citations
- General Exploratory Bonus for Optimistic Exploration in RLHFWendi Li, Changdae Oh, Sharon LiICLR 2026 · 3 citations
