What Makes Pre-trained Language Models Better Zero-shot Learners?
Jinghui Lu, Dongsheng Zhu, Weidong Han, Rui Zhao, Brian Mac Namee, Fei Tan
Abstract
Current methods for prompt learning in zeroshot scenarios widely rely on a development set with sufficient human-annotated data to select the best-performing prompt template a posteriori. This is not ideal because in a real-world zero-shot scenario of practical relevance, no labelled data is available. Thus, we propose a simple yet effective method for screening reasonable prompt templates in zero-shot text classification: Perplexity Selection (Perplection). We hypothesize that language discrepancy can be used to measure the efficacy of prompt templates, and thereby develop a substantiated perplexity-based scheme allowing for forecasting the performance of prompt templates in advance. Experiments show that our method leads to improved prediction performance in a realistic zero-shot setting, eliminating the need for any labelled examples. Dataset 1. [very/not] pleased. 2. [very/not] good. 3. [extremely/less] pleased. 4. [yellow/green] black. PPL Acc.(%) PPL Acc.(%) PPL Acc.(%) PPL Acc.(%)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ef1de4e-7d71-4553-9e51-dac5808e5a48Cited by top-tier papers9
- PaDeLLM-NER: Parallel Decoding in Large Language Models for Named Entity RecognitionJinghui Lu, Yanjie Wang, Ziwei Yang, Xuejing Liu et al.NeurIPS 2024 · 22 citations
- MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced PolicyShaoxiong Zhan, Yanlin Lai, Ziyu Lu, Dahua Lin et al.AAAI 2026 · 17 citations
- Transformers Get Stable: An End-to-End Signal Propagation Theory for Language ModelsAkhil Kedia, Mohd Abbas Zaidi, Sushil Khyalia, Jungho Jung et al.ICML 2024 · 16 citations
- MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR AdvancementWeitao Jia, Jinghui Lu, Haiyang Yu, Siqi Wang et al.AAAI 2026 · 12 citations
- Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning PerformanceYao Wang, Di Liang, Minlong PengEMNLP 2025 · 3 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
Related papers
- Pre-trained Language Models Can be Fully Zero-Shot LearnersXuandong Zhao, Siqi Ouyang, Zhiguo Yu, Ming Wu et al.ACL 2023 · 22 citations
- Interactive and Visual Prompt Engineering for Ad-hoc Task Adaptation with Large Language ModelsHendrik Strobelt, Albert Webson, Victor Sanh, Benjamin Hoover et al.IEEE VIS 2022 · 191 citations
- MetricPrompt: Prompting Model as a Relevance Metric for Few-shot Text ClassificationHongyuan Dong, Weinan Zhang, Wanxiang CheKDD 2023 · 2 citations
- A Simple Zero-shot Prompt Weighting Technique to Improve Prompt Ensembling in Text-Image ModelsJames Urquhart Allingham, Jie Ren, Michael W. Dusenberry, Xiuye Gu et al.ICML 2023 · 63 citations
- True Few-Shot Learning with Language ModelsEthan Perez, Douwe Kiela, Kyunghyun ChoNeurIPS 2021 · 547 citations
