Interpretable Deep Clustering for Tabular Data
Jonathan Svirsky, Ofir Lindenbaum
Abstract
Clustering is a fundamental learning task widely used as a first step in data analysis. For example, biologists use cluster assignments to analyze genome sequences, medical records, or images. Since downstream analysis is typically performed at the cluster level, practitioners seek reliable and interpretable clustering models. We propose a new deep-learning framework for general domain tabular data that predicts interpretable cluster assignments at the instance and cluster levels. First, we present a self-supervised procedure to identify the subset of the most informative features from each data point. Then, we design a model that predicts cluster assignments and a gate matrix that provides cluster-level feature selection. Overall, our model provides cluster assignments with an indication of the driving feature for each sample and each cluster. We show that the proposed method can reliably predict cluster assignments in biological, text, image, and physics tabular datasets. Furthermore, using previously proposed metrics, we verify that our model leads to interpretable results at a sample and cluster level. Our code is available at https://github.com/jsvir/idc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0862b3c8-6f92-43a4-abbc-8e6bb6a4b088Cited by top-tier papers7
- ZEUS: Zero-shot Embeddings for Unsupervised Separation of Tabular DataPatryk Marszalek, Tomasz Kusmierczyk, Witold Wydmanski, Jacek Tabor et al.NeurIPS 2025 · 7 citations
- Hybrid Autoencoders for Tabular Data: Leveraging Model-Based Augmentation in Low-Label SettingsErel Naor, Ofir LindenbaumNeurIPS 2025 · 6 citations
- ICR-RL: Deep Reinforcement Learning via In-Context-RegressionDavid Schiff, Ofir Lindenbaum, Yonathan EfroniICML 2026 · 5 citations
- Distributionally Robust Feature SelectionMaitreyi Swaroop, Tamar Krishnamurti, Bryan WilderNeurIPS 2025 · 2 citations
- NeuralCohort: Cohort-aware Neural Representation Learning for Healthcare AnalyticsChangshuo Liu, Lingze Zeng, Kaiping Zheng, Shaofeng Cai et al.ICML 2025
Builds on13
- Learning Diverse and Discriminative Representations via the Principle of Maximal Coding Rate ReductionYaodong Yu, Kwan Ho Ryan Chan, Chong You, Chaobing Song et al.NeurIPS 2020 · 265 citations
- Frequency Bias in Neural Networks for Input of Non-Uniform DensityRonen Basri, Meirav Galun, Amnon Geifman, David W. Jacobs et al.ICML 2020 · 229 citations
- Efficient Deep Embedded Subspace ClusteringJinyu Cai, Jicong Fan, Wenzhong Guo, Shiping Wang et al.CVPR 2022 · 127 citations
- You Never Cluster AloneYuming Shen, Ziyi Shen, Menghan Wang, Jie Qin et al.NeurIPS 2021 · 69 citations
- Locally Sparse Neural Networks for Tabular Biomedical DataJunchen Yang, Ofir Lindenbaum, Yuval KlugerICML 2022 · 45 citations
Related papers
- Deep Clustering based on Bi-Space Association LearningHao Huang, Shinjae Yoo, Chenxiao XuACM MM 2021
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- Self-Supervision Enhanced Feature Selection with Correlated GatesChanghee Lee, Fergus Imrie, Mihaela van der SchaarICLR 2022 · 26 citations
- Understanding Distributed Representations of Concepts in Deep Neural Networks without SupervisionWonjoon Chang, Dahee Kwon, Jaesik ChoiAAAI 2024 · 2 citations
- LRSC: Learning Representations for Subspace ClusteringChangsheng Li, Chen Yang, Bo Liu, Ye Yuan et al.AAAI 2021 · 16 citations
