Scalable Multi-Action Offline Policy Learning with an m-ary Tree
Shusei Eshima
摘要
Decision trees are widely used as interpretable policies for personalized treatment assignment. However, existing methods face practical challenges: binary trees can be too restrictive to capture complex heterogeneity, and tree-search algorithms often fail to scale to industrial datasets with tens of millions of observations. To address these gaps, we propose a scalable offline policy-learning method for multi-action settings. We optimize exactly over fixed-depth, bounded-branching m-ary policy trees on categorical or discretized features by maximizing an estimated policy value based on doubly robust scores. Multiway splits yield richer partitions at a given depth, enabling more expressive decision rules without deeper trees. We combine pruning, caching, and parallelization to achieve better scalability than existing algorithms. For statistical efficiency, we use nested folds to leverage all observations. For each outer fold, we hold it out for honest evaluation and further split the remaining folds into inner folds for doubly robust score computation and policy learning; we then pool results across all outer folds. Across 20 industry-scale experiments, m-ary policy trees outperform a commonly used global-winner-for-all policy in seven cases after false discovery rate control. In contrast, globally optimized binary trees do not achieve comparable gains, even at a greater depth.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Interpretable Off-Policy Learning via Hyperbox SearchDaniel Tschernutter, Tobias Hatt, Stefan FeuerriegelICML 2022 · 被引用 7 次
- Learning Prescriptive ReLU NetworksWei Sun, Asterios TsiourvasICML 2023 · 被引用 3 次
- Constrained Prescriptive Trees via Column GenerationShivaram Subramanian, Wei Sun, Youssef Drissi, Markus EttlAAAI 2022 · 被引用 12 次
- POETREE: Interpretable Policy Learning with Adaptive Decision TreesAlizée Pace, Alex J. Chan, Mihaela van der SchaarICLR 2022 · 被引用 18 次
- SPOT: Scalable Policy Optimization with Trees for Markov Decision ProcessesXuyuan Xiong, Pedro Chumpitaz-Flores, Kaixun Hua, Cheng HuaNeurIPS 2025
