A Universal Discriminator for Zero-Shot Generalization
Haike Xu, Zongyu Lin, Jing Zhou, Yanan Zheng, Zhilin Yang
Abstract
Generative modeling has been the dominant approach for large-scale pretraining and zero-shot generalization. In this work, we challenge this convention by showing that discriminative approaches perform substantially better than generative ones on a large number of NLP tasks. Technically, we train a single discriminator to predict whether a text sample comes from the true data distribution, similar to GANs. Since many NLP tasks can be formulated as selecting from a few options, we use this discriminator to predict the concatenation of input and which option has the highest probability of coming from the true data distribution. This simple formulation achieves state-of-the-art zero-shot results on the T0 benchmark, outperforming T0 by 16.0%, 7.8%, and 11.5% respectively on different scales. In the finetuning setting, our approach also achieves new state-of-the-art results on a wide range of NLP tasks, with only 1/4 parameters of previous methods. Meanwhile, our approach requires minimal prompting efforts, which largely improves robustness and is essential for real-world applications. Furthermore, we also jointly train a generalized UD in combination with generative tasks, which maintains its advantage on discriminative tasks and simultaneously works on generative tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 718b9f76-a825-4af2-9780-5229580abe9cCited by top-tier papers4
- The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-TuningSeungone Kim, Se June Joo, Doyoung Kim, Joel Jang et al.EMNLP 2023 · 45 citations
- MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction FollowingRenze Lou, Kai Zhang, Jian Xie, Yuxuan Sun et al.ICLR 2024 · 39 citations
- QLASS: Boosting Language Agent Inference via Q-Guided Stepwise SearchZongyu Lin, Yao Tang, Xingcheng Yao, Da Yin et al.ICML 2025
- Contradiction Retrieval via Contrastive Learning with SparsityHaike Xu, Zongyu Lin, Kai-Wei Chang, Yizhou Sun et al.ICML 2025
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
Related papers
- Gen-Z: Generative Zero-Shot Text Classification with Contextualized Label DescriptionsSachin Kumar, Chan Young Park, Yulia TsvetkovICLR 2024 · 8 citations
- Model-Generated Pretraining Signals Improves Zero-Shot Generalization of Text-to-Text TransformersLinyuan Gong, Chenyan Xiong, Xiaodong Liu, Payal Bajaj et al.ACL 2023
- Not All Tasks Are Born Equal: Understanding Zero-Shot GeneralizationJing Zhou, Zongyu Lin, Yanan Zheng, Jian Li et al.ICLR 2023
- Pre-trained Language Models Can be Fully Zero-Shot LearnersXuandong Zhao, Siqi Ouyang, Zhiguo Yu, Ming Wu et al.ACL 2023 · 22 citations
- Generating Training Data with Language Models: Towards Zero-Shot Language UnderstandingYu Meng, Jiaxin Huang, Yu Zhang, Jiawei HanNeurIPS 2022 · 309 citations
