A Universal Discriminator for Zero-Shot Generalization
Haike Xu, Zongyu Lin, Jing Zhou, Yanan Zheng, Zhilin Yang
摘要
Generative modeling has been the dominant approach for large-scale pretraining and zero-shot generalization. In this work, we challenge this convention by showing that discriminative approaches perform substantially better than generative ones on a large number of NLP tasks. Technically, we train a single discriminator to predict whether a text sample comes from the true data distribution, similar to GANs. Since many NLP tasks can be formulated as selecting from a few options, we use this discriminator to predict the concatenation of input and which option has the highest probability of coming from the true data distribution. This simple formulation achieves state-of-the-art zero-shot results on the T0 benchmark, outperforming T0 by 16.0%, 7.8%, and 11.5% respectively on different scales. In the finetuning setting, our approach also achieves new state-of-the-art results on a wide range of NLP tasks, with only 1/4 parameters of previous methods. Meanwhile, our approach requires minimal prompting efforts, which largely improves robustness and is essential for real-world applications. Furthermore, we also jointly train a generalized UD in combination with generative tasks, which maintains its advantage on discriminative tasks and simultaneously works on generative tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-TuningSeungone Kim, Se June Joo, Doyoung Kim, Joel Jang 等EMNLP 2023 · 被引用 45 次
- MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction FollowingRenze Lou, Kai Zhang, Jian Xie, Yuxuan Sun 等ICLR 2024 · 被引用 39 次
- QLASS: Boosting Language Agent Inference via Q-Guided Stepwise SearchZongyu Lin, Yao Tang, Xingcheng Yao, Da Yin 等ICML 2025
- Contradiction Retrieval via Contrastive Learning with SparsityHaike Xu, Zongyu Lin, Kai-Wei Chang, Yizhou Sun 等ICML 2025
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
相关 Paper
- Gen-Z: Generative Zero-Shot Text Classification with Contextualized Label DescriptionsSachin Kumar, Chan Young Park, Yulia TsvetkovICLR 2024 · 被引用 8 次
- Model-Generated Pretraining Signals Improves Zero-Shot Generalization of Text-to-Text TransformersLinyuan Gong, Chenyan Xiong, Xiaodong Liu, Payal Bajaj 等ACL 2023
- Not All Tasks Are Born Equal: Understanding Zero-Shot GeneralizationJing Zhou, Zongyu Lin, Yanan Zheng, Jian Li 等ICLR 2023
- Pre-trained Language Models Can be Fully Zero-Shot LearnersXuandong Zhao, Siqi Ouyang, Zhiguo Yu, Ming Wu 等ACL 2023 · 被引用 22 次
- Generating Training Data with Language Models: Towards Zero-Shot Language UnderstandingYu Meng, Jiaxin Huang, Yu Zhang, Jiawei HanNeurIPS 2022 · 被引用 309 次
