Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, Pontus Stenetorp
Abstract
When primed with only a handful of training samples, very large, pretrained language models such as GPT-3 have shown competitive results when compared to fully-supervised, finetuned, large, pretrained language models. We demonstrate that the order in which the samples are provided can make the difference between near state-of-the-art and random guess performance: essentially some permutations are "fantastic" and some not. We analyse this phenomenon in detail, establishing that: it is present across model sizes (even for the largest current models), it is not related to a specific subset of samples, and that a given good permutation for one model is not transferable to another. While one could use a development set to determine which permutations are performant, this would deviate from the true fewshot setting as it requires additional annotated data. Instead, we use the generative nature of language models to construct an artificial development set and based on entropy statistics of the candidate permutations on this set, we identify performant prompts. Our method yields a 13% relative improvement for GPTfamily models across eleven different established text classification tasks. 1 We can also refer these models as GPT2-base, GPT2medium, GPT2-Large, and GPT2-XL. 2 The smallest model in our experiment is the same size as BERT-base.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 61e0be47-895f-4621-9c68-722e1cbc0538Cited by top-tier papers385
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Implicit In-context LearningZhuowei Li, Zihao Xu, Ligong Han, Yunhe Gao et al.ICLR 2025 · 1,989 citations
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 887 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- True Few-Shot Learning with Language ModelsEthan Perez, Douwe Kiela, Kyunghyun ChoNeurIPS 2021 · 547 citations
- Making Pre-trained Language Models Better Few-shot LearnersTianyu Gao, Adam Fisch, Danqi ChenACL 2021
Related papers
- Cold-Start Data Selection for Better Few-shot Language Model Fine-tuning: A Prompt-based Uncertainty Propagation ApproachYue Yu, Rongzhi Zhang, Ran Xu, Jieyu Zhang et al.ACL 2023 · 12 citations
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger et al.NeurIPS 2023 · 63 citations
- Benchmarking Large Language Model Capabilities for Conditional GenerationJoshua Maynez, Priyanka Agrawal, Sebastian GehrmannACL 2023 · 6 citations
- How General-Purpose Is a Language Model? Usefulness and Safety with Human Prompters in the WildPablo Antonio Moreno Casares, Bao Sheng Loe, John Burden, Seán Ó hÉigeartaigh et al.AAAI 2022 · 3 citations
- Order-Independence Without Fine TuningReid McIlroy-Young, Katrina Brown, Conlan Olson, Linjun Zhang et al.NeurIPS 2024 · 10 citations
