PromptMix: A Class Boundary Augmentation Method for Large Language Model Distillation
Gaurav Sahu, Olga Vechtomova, Dzmitry Bahdanau, Issam H. Laradji
Abstract
Data augmentation is a widely used technique to address the problem of text classification when there is a limited amount of training data. Recent work often tackles this problem using large language models (LLMs) like GPT3 that can generate new examples given already available ones. In this work, we propose a method to generate more helpful augmented data by utilizing the LLM's abilities to follow instructions and perform few-shot classifications. Our specific PromptMix method consists of two steps: 1) generate challenging text augmentations near class boundaries; however, generating borderline examples increases the risk of false positives in the dataset, so we 2) relabel the text augmentations using a prompting-based LLM classifier to enhance the correctness of labels in the generated data. We evaluate the proposed method in challenging 2-shot and zero-shot settings on four text classification datasets: Bank-ing77, TREC6, Subjectivity (SUBJ), and Twitter Complaints. Our experiments show that generating and, crucially, relabeling borderline examples facilitates the transfer of knowledge of a massive LLM like GPT3.5-turbo into smaller and cheaper classifiers like DistilBERT base and BERT base . Furthermore, 2-shot Prompt-Mix outperforms multiple 5-shot data augmentation methods on the four datasets. Our code is available at https://github.com/ ServiceNow/PromptMix-EMNLP-2023 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd9f72b1-8cf9-4ce4-b5c6-e69c759fcaabCited by top-tier papers4
- The Efficiency vs. Accuracy Trade-off: Optimizing RAG-Enhanced LLM Recommender Systems Using Multi-Head Early ExitHuixue Zhou, Hengrui Gu, Zaifu Zhan, Xi Liu et al.ACL 2025 · 8 citations
- ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract DescriptionsSreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Chandra Kiran Reddy Evuru et al.ACL 2024 · 3 citations
- Fill In The Gaps: Model Calibration and Generalization with Synthetic DataYang Ba, Michelle Mancenido, Rong PanEMNLP 2024 · 2 citations
- Improving Clustering with Positive Pairs Generated from LLM-Driven LabelsXiaotong Zhang, Ying LiEMNLP 2025
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Do Not Have Enough Data? Deep Learning to the Rescue!Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor et al.AAAI 2020 · 398 citations
- FLEX: Unifying Evaluation for Few-Shot NLPJonathan Bragg, Arman Cohan, Kyle Lo, Iz BeltagyNeurIPS 2021 · 114 citations
- MixKD: Towards Efficient Distillation of Large-scale Language ModelsKevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou et al.ICLR 2021 · 90 citations
Related papers
- Multi-Mask Label Mapping for Prompt-Based LearningJirui Qi, Richong Zhang, Jaein Kim, Junfan Chen et al.AAAI 2023 · 1 citation
- Liberating Seen Classes: Boosting Few-Shot and Zero-Shot Text Classification via Anchor Generation and Classification ReframingHan Liu, Siyang Zhao, Xiaotong Zhang, Feng Zhang et al.AAAI 2024 · 7 citations
- inversedMixup: Data Augmentation via Inverting Mixed EmbeddingsFanshuang Kong, Richong Zhang, Qiyu Sun, Zhijie Nie et al.KDD 2026
- Incubating Text Classifiers Following User Instruction with Nothing but LLMLetian Peng, Zilong Wang, Jingbo ShangEMNLP 2024
- Language Models are Weak LearnersHariharan Manikandan, Yiding Jiang, J. Zico KolterNeurIPS 2023 · 32 citations
