Subgroup Discovery with the Cox Model
Zachary Izzo, Iain Melvin
Abstract
We study the problem of subgroup discovery for survival analysis, where the goal is to find an interpretable subset of the data on which a Cox model is highly accurate. Our work is the first to study this particular subgroup problem, for which we make several contributions. Subgroup discovery methods generally require a "quality function" in order to sift through and select the most advantageous subgroups. We first examine why existing natural choices for quality functions are insufficient to solve the subgroup discovery problem for the Cox model. To address the shortcomings of existing metrics, we introduce two technical innovations: the expected prediction entropy (EPE), a novel metric for evaluating survival models which predict a hazard function; and the conditional rank statistics (CRS), a statistical object which quantifies the deviation of an individual point to the distribution of survival times in an existing subgroup. We study the EPE and CRS theoretically and show that they can solve many of the problems with existing metrics. Having established the fundamentals of the problem, we then turn to methodology for solving it. To this end, we introduce a total of eight algorithms for the Cox subgroup discovery problem. The main algorithm, which is based on the DDGroup framework of [22] , is able to take advantage of both the EPE and the CRS, allowing us to give theoretical correctness results for this algorithm in a well-specified setting. We evaluate all of the proposed methods empirically on both synthetic and real data. The experiments confirm our theory, showing that our contributions allow for the recovery of a ground-truth subgroup in well-specified cases, as well as leading to better model fit compared to naively fitting the Cox model to the whole dataset in practical settings. Lastly, to showcase the utility of the subgroups themselves beyond improving model predictions, we conduct a case study on jet engine simulation data from NASA. The discovered subgroups uncover known nonlinearities/homogeneity in the data, and which suggest design choices which have been mirrored in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- An Effective Meaningful Way to Evaluate Survival ModelsShiang Qi, Neeraj Kumar, Mahtab Farrokh, Weijie Sun et al.ICML 2023 · 28 citations
- Proper Scoring Rules for Survival AnalysisHiroki YanagisawaICML 2023 · 12 citations
- Neural Frailty Machine: Beyond proportional hazard assumption in neural survival regressionsRuofan Wu, Jiawei Qiao, Mingzhe Wu, Wen Yu et al.NeurIPS 2023 · 9 citations
- Learning Exceptional Subgroups by End-to-End Maximizing KL-DivergenceSascha Xu, Nils Philipp Walter, Janis Kalofolias, Jilles VreekenICML 2024 · 8 citations
- Data-Driven Subgroup Identification for Linear RegressionZachary Izzo, Ruishan Liu, James ZouICML 2023 · 7 citations
Related papers
- Optimal Survival Trees: A Dynamic Programming ApproachTim Huisman, Jacobus G. M. van der Linden, Emir DemirovicAAAI 2024 · 8 citations
- REDS: Rule Extraction for Discovering ScenariosVadim Arzamasov, Klemens BöhmSIGMOD 2021 · 2 citations
- FastSurvival: Hidden Computational Blessings in Training Cox Proportional Hazards ModelsJiachang Liu, Rui Zhang, Cynthia RudinNeurIPS 2024 · 1 citation
- KSP: Kolmogorov-Smirnov metric-based Post-Hoc Calibration for Survival AnalysisJeongho Park, Daheen Kim, Cheoljun Kim, Hyungbin Park et al.NeurIPS 2025 · 2 citations
- Functional Decomposition and Shapley Interactions for Interpreting Survival ModelsSophie Hanna Langbein, Hubert Baniecki, Fabian Fumagalli, Niklas Koenen et al.ICML 2026
