Lune

KDD2026顶会

Scalable Multi-Action Offline Policy Learning with an m-ary Tree

Shusei Eshima

2026年份

摘要

Decision trees are widely used as interpretable policies for personalized treatment assignment. However, existing methods face practical challenges: binary trees can be too restrictive to capture complex heterogeneity, and tree-search algorithms often fail to scale to industrial datasets with tens of millions of observations. To address these gaps, we propose a scalable offline policy-learning method for multi-action settings. We optimize exactly over fixed-depth, bounded-branching m-ary policy trees on categorical or discretized features by maximizing an estimated policy value based on doubly robust scores. Multiway splits yield richer partitions at a given depth, enabling more expressive decision rules without deeper trees. We combine pruning, caching, and parallelization to achieve better scalability than existing algorithms. For statistical efficiency, we use nested folds to leverage all observations. For each outer fold, we hold it out for honest evaluation and further split the remaining folds into inner folds for doubly robust score computation and policy learning; we then pool results across all outer folds. Across 20 industry-scale experiments, m-ary policy trees outperform a commonly used global-winner-for-all policy in seven cases after false discovery rate control. In contrast, globally optimized binary trees do not achieve comparable gains, even at a greater depth.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖