C2TC: A Training-Free Framework for Efficient Tabular Data Condensation
Sijia Xu, Fan Li, Xiaoyang Wang, Zhengyi Yang, Xuemin Lin
Abstract
Tabular data, organized in rows and columns, represents the most common data format in industrial relational databases, underpinning modern data analytics and decisionmaking. However, the ever-increasing scale of tabular data poses significant computational and storage challenges to learningbased analytical systems. This highlights the need for dataefficient learning, which maximizes the utility of available data to enable effective model training and generalization using substantially fewer samples. Dataset condensation (DC) has recently emerged as a promising data-centric paradigm that synthesizes small yet informative datasets to preserve data utility while greatly reducing storage and training costs. However, existing DC methods are computationally intensive due to reliance on complex gradient-based optimization. Moreover, they often overlook key characteristics of tabular data, such as heterogeneous features and class imbalance. To address these limitations, we introduce (Class-Adaptive Clustering for Tabular Condensation), the first training-free tabular dataset condensation framework that jointly optimizes class allocation and feature representation, enabling efficient and scalable condensation. Specifically, we reformulate the dataset condensation objective into a novel class-adaptive cluster allocation problem (CCAP), which eliminates costly training and integrates adaptive label allocation to handle class imbalance. To solve the NP-hard CCAP, we develop HFILS, a heuristic local search that alternates between soft allocation and class-wise clustering to efficiently obtain high-quality solutions. Moreover, a hybrid categorical feature encoding (HCFE) is proposed for semantics-preserving clustering of heterogeneous discrete attributes. Extensive experiments on 10 real-world datasets demonstrate that improves efficiency by at least 2 orders of magnitude over state-of-the-art baselines, while achieving superior downstream performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Graph is a Natural Regularization: Revisiting Vector Quantization for Graph Representation LearningZian Zhai, Fan Li, Xingyu Tan, Xiaoyang Wang et al.ICML 2026 · 2 citations
- Anchor-guided Hypergraph Condensation with Dual-level DiscriminationFan Li, Xiaoyang Wang, Chen Chen, Wenjie ZhangICML 2026
- Let the Prototype Guide You: Robust Aggregation of Sparse Multi-Class Annotations via Annotator Prototype LearningJu Chen, Jun Feng, Shenyu ZhangICML 2026
Builds on27
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 2,148 citations
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 684 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
Related papers
- Rethinking and Accelerating Graph Condensation: A Training-Free Approach with Class PartitionXinyi Gao, Guanhua Ye, Tong Chen, Wentao Zhang et al.WWW 2025 · 27 citations
- Elucidating the Design Space of Dataset CondensationShitong Shao, Zikai Zhou, Huanran Chen, Zhiqiang ShenNeurIPS 2024 · 47 citations
- Privacy for Free: How does Dataset Condensation Help Privacy?Tian Dong, Bo Zhao, Lingjuan LyuICML 2022 · 154 citations
- Dataset Condensation via Efficient Synthetic-Data ParameterizationJang-Hyun Kim, Jinuk Kim, Seong Joon Oh, Sangdoo Yun et al.ICML 2022 · 234 citations
- Improved Distribution Matching for Dataset CondensationGanlong Zhao, Guanbin Li, Yipeng Qin, Yizhou YuCVPR 2023
