Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even Performance
Branislav Pecher, Ivan Srba, Mária Bieliková
Abstract
When solving NLP tasks with limited labelled data, researchers typically either use a general large language model without further update, or use a small number of labelled samples to tune a specialised smaller model. In this work, we answer an important question -how many labelled samples are required for the specialised small models to outperform general large models, while taking the performance variance into consideration. By observing the behaviour of fine-tuning, instruction-tuning, prompting and in-context learning on 8 language models, we identify such performance break-even points across 8 representative text classification tasks of varying characteristics. We show that the specialised models often need only few samples (on average 100) to be on par or better than the general ones. At the same time, the number of required labels strongly depends on the dataset or task characteristics, with fine-tuning on binary datasets requiring significantly more samples. When performance variance is taken into consideration, the number of required labels increases on average by 100 -200%. Finally, larger models do not consistently lead to better performance and lower variance, with 4-bit quantisation having negligible impact.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a8b7f89-9b9b-405d-8703-a48724d911f3Cited by top-tier papers4
- The Representation Landscape of Few-Shot Learning and Fine-Tuning in Large Language ModelsDiego Doimo, Alessandro Serra, Alessio Ansuini, Alberto CazzanigaNeurIPS 2024 · 21 citations
- Structuring Radiology Reports: Challenging LLMs with Lightweight ModelsJohannes Moll, Louisa Fay, Asfandyar Azhar, Sophie Ostmeier et al.EMNLP 2025 · 4 citations
- From Weights to Activations: Is Steering the Next Frontier of Adaptation?Simon Ostermann, Daniil Gurgurov, Tanja Baeumel, Michael A. Hedderich et al.ACL 2026 · 3 citations
- Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning PerspectiveBishwamittra Ghosh, Soumi Das, Till Speicher, Qinyuan Wu et al.ACL 2026
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
Related papers
- Two-stage LLM Fine-tuning with Less Specialization and More GeneralizationYihan Wang, Si Si, Daliang Li, Michal Lukasik et al.ICLR 2024 · 45 citations
- Meta-learning via Language Model In-context TuningYanda Chen, Ruiqi Zhong, Sheng Zha, George Karypis et al.ACL 2022
- Evaluating the Zero-shot Robustness of Instruction-tuned Language ModelsJiuding Sun, Chantal Shaib, Byron C. WallaceICLR 2024 · 75 citations
- RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data SelectionYixin Yang, Qingxiu Dong, Linli Yao, Fangwei Zhu et al.AAAI 2026
- Specialist or Generalist? Instruction Tuning for Specific NLP TasksChufan Shi, Yixuan Su, Cheng Yang, Yujiu Yang et al.EMNLP 2023 · 11 citations
