Interpretable Off-Policy Learning via Hyperbox Search
Daniel Tschernutter, Tobias Hatt, Stefan Feuerriegel
Abstract
Personalized treatment decisions have become an integral part of modern medicine. Thereby, the aim is to make treatment decisions based on individual patient characteristics. Numerous methods have been developed for learning such policies from observational data that achieve the best outcome across a certain policy class. Yet these methods are rarely interpretable. However, interpretability is often a prerequisite for policy learning in clinical practice. In this paper, we propose an algorithm for interpretable off-policy learning via hyperbox search. In particular, our policies can be represented in disjunctive normal form (i.e., OR-of-ANDs) and are thus intelligible. We prove a universal approximation theorem that shows that our policy class is flexible enough to approximate any measurable function arbitrarily well. For optimization, we develop a tailored column generation procedure within a branch-and-bound framework. Using a simulation study, we demonstrate that our algorithm outperforms state-of-the-art methods from interpretable off-policy learning in terms of regret. Using real-word clinical data, we perform a user study with actual clinical experts, who rate our policies as highly interpretable.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Reliable Off-Policy Learning for Dosage CombinationsJonas Schweisthal, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelNeurIPS 2023 · 22 citations
- Treatment Effect Estimation for Optimal Decision-MakingDennis Frauen, Valentyn Melnychuk, Jonas Schweisthal, Mihaela van der Schaar et al.NeurIPS 2025 · 8 citations
- Efficient and Sharp Off-Policy Learning under Unobserved ConfoundingKonstantin Hess, Dennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICLR 2026 · 5 citations
- AIRS: Explanation for Deep Reinforcement Learning based Security ApplicationsJiahao Yu, Wenbo Guo, Qi Qin, Gang Wang et al.USENIX Security 2023
Builds on1
Related papers
- Learning Prescriptive ReLU NetworksWei Sun, Asterios TsiourvasICML 2023 · 3 citations
- Scalable Multi-Action Offline Policy Learning with an m-ary TreeShusei EshimaKDD 2026
- Logic-Logit: A Logic-Based Approach to Choice ModelingShuhan Zhang, Wendi Ren, Shuang LiICLR 2025
- Model Distillation for Revenue Optimization: Interpretable Personalized PricingMax Biggs, Wei Sun, Markus EttlICML 2021 · 42 citations
- POETREE: Interpretable Policy Learning with Adaptive Decision TreesAlizée Pace, Alex J. Chan, Mihaela van der SchaarICLR 2022 · 18 citations
