C2TC: A Training-Free Framework for Efficient Tabular Data Condensation
Sijia Xu, Fan Li, Xiaoyang Wang, Zhengyi Yang, Xuemin Lin
摘要
Tabular data, organized in rows and columns, represents the most common data format in industrial relational databases, underpinning modern data analytics and decisionmaking. However, the ever-increasing scale of tabular data poses significant computational and storage challenges to learningbased analytical systems. This highlights the need for dataefficient learning, which maximizes the utility of available data to enable effective model training and generalization using substantially fewer samples. Dataset condensation (DC) has recently emerged as a promising data-centric paradigm that synthesizes small yet informative datasets to preserve data utility while greatly reducing storage and training costs. However, existing DC methods are computationally intensive due to reliance on complex gradient-based optimization. Moreover, they often overlook key characteristics of tabular data, such as heterogeneous features and class imbalance. To address these limitations, we introduce (Class-Adaptive Clustering for Tabular Condensation), the first training-free tabular dataset condensation framework that jointly optimizes class allocation and feature representation, enabling efficient and scalable condensation. Specifically, we reformulate the dataset condensation objective into a novel class-adaptive cluster allocation problem (CCAP), which eliminates costly training and integrates adaptive label allocation to handle class imbalance. To solve the NP-hard CCAP, we develop HFILS, a heuristic local search that alternates between soft allocation and class-wise clustering to efficiently obtain high-quality solutions. Moreover, a hybrid categorical feature encoding (HCFE) is proposed for semantics-preserving clustering of heterogeneous discrete attributes. Extensive experiments on 10 real-world datasets demonstrate that improves efficiency by at least 2 orders of magnitude over state-of-the-art baselines, while achieving superior downstream performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Graph is a Natural Regularization: Revisiting Vector Quantization for Graph Representation LearningZian Zhai, Fan Li, Xingyu Tan, Xiaoyang Wang 等ICML 2026 · 被引用 2 次
- Anchor-guided Hypergraph Condensation with Dual-level DiscriminationFan Li, Xiaoyang Wang, Chen Chen, Wenjie ZhangICML 2026
- Let the Prototype Guide You: Robust Aggregation of Sparse Multi-Class Annotations via Annotator Prototype LearningJu Chen, Jun Feng, Shenyu ZhangICML 2026
它引用的顶会 Paper27
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- TabNet: Attentive Interpretable Tabular LearningSercan Ö. Arik, Tomas PfisterAAAI 2021 · 被引用 2,148 次
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 被引用 1,847 次
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
相关 Paper
- Rethinking and Accelerating Graph Condensation: A Training-Free Approach with Class PartitionXinyi Gao, Guanhua Ye, Tong Chen, Wentao Zhang 等WWW 2025 · 被引用 27 次
- Elucidating the Design Space of Dataset CondensationShitong Shao, Zikai Zhou, Huanran Chen, Zhiqiang ShenNeurIPS 2024 · 被引用 47 次
- Privacy for Free: How does Dataset Condensation Help Privacy?Tian Dong, Bo Zhao, Lingjuan LyuICML 2022 · 被引用 154 次
- Dataset Condensation via Efficient Synthetic-Data ParameterizationJang-Hyun Kim, Jinuk Kim, Seong Joon Oh, Sangdoo Yun 等ICML 2022 · 被引用 234 次
- Improved Distribution Matching for Dataset CondensationGanlong Zhao, Guanbin Li, Yipeng Qin, Yizhou YuCVPR 2023
