Interpretable Attribute Discretization
Eugenie Lai, Inbal Croitoru, Brit Youngmann, Sainyam Galhotra, El Kindi Rezig, Michael J. Cafarella
摘要
Data discretization, the conversion of numeric attributes into categorical bins, underpins many data-driven tasks, from visualization and causal analysis to symbolic learning. Interpretable partitions make results easier to communicate and act upon. Yet, a fundamental tension underlies this discretization process: partitions that are interpretable to humans often differ from those that maximize utility performance (e.g., causal analysis). Despite the abundance of discretization techniques, reconciling this utility-semantics trade-off remains an open challenge. In this work, we formally define the interpretable attribute discretization problem and introduce a reinforcement-learning-based framework that jointly optimizes utility and semantic fidelity. Our approach learns a partition selection policy that can navigate a vast space of candidates and accurately estimate the optimal partitions by strategically sampling only 200 candidates out of millions. We measure success by the average Hausdorff distance between the estimated and true Pareto frontier of the utility–semantics trade-off. On seven datasets across four tasks (visualization, modeling, imputation, causal analysis) and multiple semantic measures, our method achieves the lowest overall error, typically 2–4 × smaller than strong alternatives, yielding a more accurate Pareto frontier that reveals the complete utility–semantics trade-off.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Weakly-Supervised Reinforcement Learning for Controllable BehaviorLisa Lee, Ben Eysenbach, Ruslan Salakhutdinov, Shixiang Shane Gu 等NeurIPS 2020 · 被引用 28 次
- DisCo-DSO: Coupling Discrete and Continuous Optimization for Efficient Generative Design in Hybrid SpacesJacob F. Pettit, Chak Shing Lee, Jiachen Yang, Alex Ho 等AAAI 2025
- Guided Exploration of Data SummariesBrit Youngmann, Sihem Amer-Yahia, Aurélien PersonnazVLDB 2022 · 被引用 22 次
- Interpretable Off-Policy Learning via Hyperbox SearchDaniel Tschernutter, Tobias Hatt, Stefan FeuerriegelICML 2022 · 被引用 7 次
- Adaptive Discrete Communication Bottlenecks with Dynamic Vector Quantization for Heterogeneous Representational CoarsenessDianbo Liu, Alex Lamb, Xu Ji, Pascal Tikeng Notsawo Jr. 等AAAI 2023 · 被引用 19 次
