Range-limited Augmentation for Few-shot Learning in Tabular Data with Comprehensive Benchmark
Kyungeun Lee, Moonjung Eo, Hye-Seung Cho, Min-Kook Suh, Seoyoon Kim, Ye Seul Sim, Suhee Yoon, Sanghyu Yoon, Woohyung Lim
Abstract
Few-shot learning is crucial for tabular data, where the high cost of annotation often limits the availability of labeled samples. Despite its importance in real-world applications such as healthcare and finance, few-shot learning in tabular domains has received limited attention. To address this, we introduce range-limited augmentation, a novel augmentation strategy for contrastive learning that perturbs numerical features within predefined feature-specific ranges. Unlike conventional augmentations, which may result in false positive pairs during contrastive learning, our approach ensures semantic consistency by restricting augmentations to ranges. A quantitative analysis confirms that range-limited augmentation better preserves task-relevant information compared to existing augmentation techniques. Additionally, we propose FeSTa (Few-Shot Tabular classification benchmark), the first large-scale benchmark designed to systematically evaluate few-shot learning methods in tabular data. FeSTa includes 50 datasets and 32 algorithms spanning supervised, unsupervised, self-supervised, semi-supervised, and foundation models. Experiments on FeSTa show that range-limited augmentation consistently ranks among the top methods, achieving an average rank of 2.6 out of 32 in 1-shot classification, despite not relying on large-scale pretraining or complex architectures. The benchmark code is available in https://github.com/kyungeun-lee/festa.git.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- D2R2: Diffusion-based Representation with Random Distance Matching for Tabular Few-shot LearningRuoxue Liu, Linjiajie Fang, Wenjia Wang, Bingyi JingNeurIPS 2024 · 7 citations
- STUNT: Few-shot Tabular Learning with Self-generated Tasks from Unlabeled TablesJaehyun Nam, Jihoon Tack, Kyungmin Lee, Hankook Lee et al.ICLR 2023 · 2 citations
- Supervised Contrastive Few-Shot Learning for High-Frequency Time SeriesXi Chen, Cheng Ge, Ming Wang, Jin WangAAAI 2023 · 15 citations
- Scarf: Self-Supervised Contrastive Learning using Random Feature CorruptionDara Bahri, Heinrich Jiang, Yi Tay, Donald MetzlerICLR 2022 · 233 citations
- ConFeSS: A Framework for Single Source Cross-Domain Few-Shot LearningDebasmit Das, Sungrack Yun, Fatih PorikliICLR 2022 · 57 citations
