Lune

KDD2026Top-tier venue

Scalable Multi-Action Offline Policy Learning with an m-ary Tree

Shusei Eshima

2026Year

Abstract

Decision trees are widely used as interpretable policies for personalized treatment assignment. However, existing methods face practical challenges: binary trees can be too restrictive to capture complex heterogeneity, and tree-search algorithms often fail to scale to industrial datasets with tens of millions of observations. To address these gaps, we propose a scalable offline policy-learning method for multi-action settings. We optimize exactly over fixed-depth, bounded-branching m-ary policy trees on categorical or discretized features by maximizing an estimated policy value based on doubly robust scores. Multiway splits yield richer partitions at a given depth, enabling more expressive decision rules without deeper trees. We combine pruning, caching, and parallelization to achieve better scalability than existing algorithms. For statistical efficiency, we use nested folds to leverage all observations. For each outer fold, we hold it out for honest evaluation and further split the remaining folds into inner folds for doubly robust score computation and policy learning; we then pool results across all outer folds. Across 20 industry-scale experiments, m-ary policy trees outperform a commonly used global-winner-for-all policy in seven cases after false discovery rate control. In contrast, globally optimized binary trees do not achieve comparable gains, even at a greater depth.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 5f62839c-bf0b-4cc9-bfff-0760cc0df2e3

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines