Integrating Reward Maximization and Population Estimation: Sequential Decision-Making for Internal Revenue Service Audit Selection
Peter Henderson, Ben Chugg, Brandon R. Anderson, Kristen M. Altenburger, Alex Turk, John Guyton, Jacob S. Goldin, Daniel E. Ho
摘要
We introduce a new setting, optimize-and-estimate structured bandits. Here, a policy must select a batch of arms, each characterized by its own context, that would allow it to both maximize reward and maintain an accurate (ideally unbiased) population estimate of the reward. This setting is inherent to many public and private sector applications and often requires handling delayed feedback, small data, and distribution shifts. We demonstrate its importance on real data from the United States Internal Revenue Service (IRS). The IRS performs yearly audits of the tax base. Two of its most important objectives are to identify suspected misreporting and to estimate the "tax gap" -- the global difference between the amount paid and true amount owed. Based on a unique collaboration with the IRS, we cast these two processes as a unified optimize-and-estimate structured bandit. We analyze optimize-and-estimate approaches to the IRS problem and propose a novel mechanism for unbiased population estimation that achieves rewards comparable to baseline approaches. This approach has the potential to improve audit efficacy, while maintaining policy-relevant estimates of the tax gap. This has important social consequences given that the current tax gap is estimated at nearly half a trillion dollars. We suggest that this problem setting is fertile ground for further research and we highlight its interesting challenges. The results of this and related research are currently being incorporated into the continual improvement of the IRS audit selection methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Auditing Fairness by BettingBen Chugg, Santiago Cortes-Gomez, Bryan Wilder, Aaditya RamdasNeurIPS 2023 · 被引用 29 次
- Domain constraints improve risk prediction when outcome data is missingSidhika Balachandar, Nikhil Garg, Emma PiersonICLR 2024 · 被引用 11 次
- Entropy Regularization for Population EstimationBen Chugg, Peter Henderson, Jacob S. Goldin, Daniel E. HoAAAI 2023 · 被引用 3 次
- From Biased Selective Labels to Pseudo-Labels: An Expectation-Maximization Framework for Learning from Biased DecisionsTrenton Chang, Jenna WiensICML 2024 · 被引用 1 次
它引用的顶会 Paper6
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Beyond UCB: Optimal and Efficient Contextual Bandits with Regression OraclesDylan J. Foster, Alexander RakhlinICML 2020 · 被引用 241 次
- On Statistical Bias In Active Learning: How and When to Fix ItSebastian Farquhar, Yarin Gal, Tom RainforthICLR 2021 · 被引用 96 次
- Contextual Bandits with Large Action Spaces: Made PracticalYinglun Zhu, Dylan J. Foster, John Langford, Paul MineiroICML 2022 · 被引用 34 次
- Top-k eXtreme Contextual Bandits with Arm HierarchyRajat Sen, Alexander Rakhlin, Lexing Ying, Rahul Kidambi 等ICML 2021 · 被引用 16 次
相关 Paper
- Cost-Effective Incentive Allocation via Structured Counterfactual InferenceRomain Lopez, Chenchen Li, Xiang Yan, Junwu Xiong 等AAAI 2020 · 被引用 21 次
- Optimally Auditing Adversarial AgentsSanmay Das, Fang-Yi Yu, Yuang ZhangAAAI 2026 · 被引用 1 次
- Inference for Batched BanditsKelly W. Zhang, Lucas Janson, Susan A. MurphyNeurIPS 2020 · 被引用 115 次
- Statistical Inference on Multi-armed Bandits with Delayed FeedbackLei Shi, Jingshen Wang, Tianhao WuICML 2023 · 被引用 7 次
- (Almost) Free Incentivized Exploration from Decentralized Learning AgentsChengshuai Shi, Haifeng Xu, Wei Xiong, Cong ShenNeurIPS 2021 · 被引用 10 次
