DEALT: LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text Classification
Wayne Lu, Xiaoxi Cui
Abstract
Real-world text classification datasets frequently exhibit long-tail distributions, where numerous classes have sparse data, significantly degrading model performance on these underrepresented categories. While Large Language Models (LLMs) offer promise for data augmentation, existing methods often produce semantically limited samples, neglect "implicit long-tails" (sparse sub-patterns within classes), and lack cost-effective optimization. To address these challenges, we propose DEALT (LLM-driven Diversity-Enhanced Data Augmentation for Long-Tail Text Classification), a novel cognitive-inspired framework emulating the human learning process of "recognize, explore, generate, and optimize." DEALT systematically enhances augmented data diversity by first detecting both explicit and implicit long-tails. It then employs an LLM for diversity-aware planning of augmentation strategies, followed by conditional generation. A lowoverhead quality and diversity validator filters the synthetic data, and an adaptive incremental sampler refines future augmentation efforts based on proxy model feedback, ensuring efficient and budget-aware optimization. Extensive experiments on multiple public text classification datasets demonstrate DEALT's superiority over state-of-the-art methods in improving tail-class performance and overall model robustness by generating more diverse and high-fidelity augmented data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c31bfcc9-d3b2-4727-975f-239f84d10b73Cited by top-tier papers3
- Leveraging Machine Unlearning for Cost-Efficient Preference AlignmentXiaoHua Feng, Yuyuan Li, HuWei Ji, Li Zhang et al.ICML 2026 · 4 citations
- Retrieval-Augmented Multimodal Model for Fake News DetectionYiheng Li, Weihai Lu, Hanyi Yu, Yue WangSIGIR 2026 · 4 citations
- MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance DetectionWeihai Lu, Zhejun Zhao, Yanshu Li, Huan HeACL 2026
Builds on14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Turning large language models into cognitive modelsMarcel Binz, Eric SchulzICLR 2024 · 99 citations
- TELEClass: Taxonomy Enrichment and LLM-Enhanced Hierarchical Text Classification with Minimal SupervisionYunyi Zhang, Ruozhen Yang, Xueqiang Xu, Rui Li et al.WWW 2025 · 53 citations
Related papers
- LTGC: Long-Tail Recognition via Leveraging LLMs-Driven Generated ContentQihao Zhao, Yalun Dai, Hao Li, Wei Hu et al.CVPR 2024 · 22 citations
- LLM-AutoDA: Large Language Model-Driven Automatic Data Augmentation for Long-tailed ProblemsPengkun Wang, Zhe Zhao, Haibin Wen, Fanfu Wang et al.NeurIPS 2024 · 26 citations
- Long-Tailed Classification with Multi-Granularity SemanticsYuting Liu, Liu Yang, Yu WangICCV 2025 · 1 citation
- Label-Specific Feature Augmentation for Long-Tailed Multi-Label Text ClassificationPengyu Xu, Lin Xiao, Bing Liu, Sijin Lu et al.AAAI 2023 · 29 citations
- Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed ClassificationXiaohua Chen, Yucan Zhou, Dayan Wu, Wanqian Zhang et al.AAAI 2022 · 66 citations
