Scalable Multi-Action Offline Policy Learning with an m-ary Tree
Shusei Eshima
Abstract
Decision trees are widely used as interpretable policies for personalized treatment assignment. However, existing methods face practical challenges: binary trees can be too restrictive to capture complex heterogeneity, and tree-search algorithms often fail to scale to industrial datasets with tens of millions of observations. To address these gaps, we propose a scalable offline policy-learning method for multi-action settings. We optimize exactly over fixed-depth, bounded-branching m-ary policy trees on categorical or discretized features by maximizing an estimated policy value based on doubly robust scores. Multiway splits yield richer partitions at a given depth, enabling more expressive decision rules without deeper trees. We combine pruning, caching, and parallelization to achieve better scalability than existing algorithms. For statistical efficiency, we use nested folds to leverage all observations. For each outer fold, we hold it out for honest evaluation and further split the remaining folds into inner folds for doubly robust score computation and policy learning; we then pool results across all outer folds. Across 20 industry-scale experiments, m-ary policy trees outperform a commonly used global-winner-for-all policy in seven cases after false discovery rate control. In contrast, globally optimized binary trees do not achieve comparable gains, even at a greater depth.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5f62839c-bf0b-4cc9-bfff-0760cc0df2e3Related papers
- Interpretable Off-Policy Learning via Hyperbox SearchDaniel Tschernutter, Tobias Hatt, Stefan FeuerriegelICML 2022 · 7 citations
- Learning Prescriptive ReLU NetworksWei Sun, Asterios TsiourvasICML 2023 · 3 citations
- Constrained Prescriptive Trees via Column GenerationShivaram Subramanian, Wei Sun, Youssef Drissi, Markus EttlAAAI 2022 · 12 citations
- POETREE: Interpretable Policy Learning with Adaptive Decision TreesAlizée Pace, Alex J. Chan, Mihaela van der SchaarICLR 2022 · 18 citations
- SPOT: Scalable Policy Optimization with Trees for Markov Decision ProcessesXuyuan Xiong, Pedro Chumpitaz-Flores, Kaixun Hua, Cheng HuaNeurIPS 2025
