Interpretable Deep Clustering for Tabular Data
Jonathan Svirsky, Ofir Lindenbaum
摘要
Clustering is a fundamental learning task widely used as a first step in data analysis. For example, biologists use cluster assignments to analyze genome sequences, medical records, or images. Since downstream analysis is typically performed at the cluster level, practitioners seek reliable and interpretable clustering models. We propose a new deep-learning framework for general domain tabular data that predicts interpretable cluster assignments at the instance and cluster levels. First, we present a self-supervised procedure to identify the subset of the most informative features from each data point. Then, we design a model that predicts cluster assignments and a gate matrix that provides cluster-level feature selection. Overall, our model provides cluster assignments with an indication of the driving feature for each sample and each cluster. We show that the proposed method can reliably predict cluster assignments in biological, text, image, and physics tabular datasets. Furthermore, using previously proposed metrics, we verify that our model leads to interpretable results at a sample and cluster level. Our code is available at https://github.com/jsvir/idc.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- ZEUS: Zero-shot Embeddings for Unsupervised Separation of Tabular DataPatryk Marszalek, Tomasz Kusmierczyk, Witold Wydmanski, Jacek Tabor 等NeurIPS 2025 · 被引用 7 次
- Hybrid Autoencoders for Tabular Data: Leveraging Model-Based Augmentation in Low-Label SettingsErel Naor, Ofir LindenbaumNeurIPS 2025 · 被引用 6 次
- ICR-RL: Deep Reinforcement Learning via In-Context-RegressionDavid Schiff, Ofir Lindenbaum, Yonathan EfroniICML 2026 · 被引用 5 次
- Distributionally Robust Feature SelectionMaitreyi Swaroop, Tamar Krishnamurti, Bryan WilderNeurIPS 2025 · 被引用 2 次
- NeuralCohort: Cohort-aware Neural Representation Learning for Healthcare AnalyticsChangshuo Liu, Lingze Zeng, Kaiping Zheng, Shaofeng Cai 等ICML 2025
它引用的顶会 Paper13
- Learning Diverse and Discriminative Representations via the Principle of Maximal Coding Rate ReductionYaodong Yu, Kwan Ho Ryan Chan, Chong You, Chaobing Song 等NeurIPS 2020 · 被引用 265 次
- Frequency Bias in Neural Networks for Input of Non-Uniform DensityRonen Basri, Meirav Galun, Amnon Geifman, David W. Jacobs 等ICML 2020 · 被引用 229 次
- Efficient Deep Embedded Subspace ClusteringJinyu Cai, Jicong Fan, Wenzhong Guo, Shiping Wang 等CVPR 2022 · 被引用 127 次
- You Never Cluster AloneYuming Shen, Ziyi Shen, Menghan Wang, Jie Qin 等NeurIPS 2021 · 被引用 69 次
- Locally Sparse Neural Networks for Tabular Biomedical DataJunchen Yang, Ofir Lindenbaum, Yuval KlugerICML 2022 · 被引用 45 次
相关 Paper
- Deep Clustering based on Bi-Space Association LearningHao Huang, Shinjae Yoo, Chenxiao XuACM MM 2021
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 被引用 2,148 次
- Self-Supervision Enhanced Feature Selection with Correlated GatesChanghee Lee, Fergus Imrie, Mihaela van der SchaarICLR 2022 · 被引用 26 次
- Understanding Distributed Representations of Concepts in Deep Neural Networks without SupervisionWonjoon Chang, Dahee Kwon, Jaesik ChoiAAAI 2024 · 被引用 2 次
- LRSC: Learning Representations for Subspace ClusteringChangsheng Li, Chen Yang, Bo Liu, Ye Yuan 等AAAI 2021 · 被引用 16 次
