Interpretable Attribute Discretization
Eugenie Lai, Inbal Croitoru, Brit Youngmann, Sainyam Galhotra, El Kindi Rezig, Michael J. Cafarella
Abstract
Data discretization, the conversion of numeric attributes into categorical bins, underpins many data-driven tasks, from visualization and causal analysis to symbolic learning. Interpretable partitions make results easier to communicate and act upon. Yet, a fundamental tension underlies this discretization process: partitions that are interpretable to humans often differ from those that maximize utility performance (e.g., causal analysis). Despite the abundance of discretization techniques, reconciling this utility-semantics trade-off remains an open challenge. In this work, we formally define the interpretable attribute discretization problem and introduce a reinforcement-learning-based framework that jointly optimizes utility and semantic fidelity. Our approach learns a partition selection policy that can navigate a vast space of candidates and accurately estimate the optimal partitions by strategically sampling only 200 candidates out of millions. We measure success by the average Hausdorff distance between the estimated and true Pareto frontier of the utility–semantics trade-off. On seven datasets across four tasks (visualization, modeling, imputation, causal analysis) and multiple semantic measures, our method achieves the lowest overall error, typically 2–4 × smaller than strong alternatives, yielding a more accurate Pareto frontier that reveals the complete utility–semantics trade-off.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6ffab991-de15-40fd-9946-69ac65c3e3b5Related papers
- Weakly-Supervised Reinforcement Learning for Controllable BehaviorLisa Lee, Ben Eysenbach, Ruslan Salakhutdinov, Shixiang Shane Gu et al.NeurIPS 2020 · 28 citations
- DisCo-DSO: Coupling Discrete and Continuous Optimization for Efficient Generative Design in Hybrid SpacesJacob F. Pettit, Chak Shing Lee, Jiachen Yang, Alex Ho et al.AAAI 2025
- Guided Exploration of Data SummariesBrit Youngmann, Sihem Amer-Yahia, Aurélien PersonnazVLDB 2022 · 22 citations
- Interpretable Off-Policy Learning via Hyperbox SearchDaniel Tschernutter, Tobias Hatt, Stefan FeuerriegelICML 2022 · 7 citations
- Adaptive Discrete Communication Bottlenecks with Dynamic Vector Quantization for Heterogeneous Representational CoarsenessDianbo Liu, Alex Lamb, Xu Ji, Pascal Tikeng Notsawo Jr. et al.AAAI 2023 · 19 citations
