Benchmarking Large Language Model Capabilities for Conditional Generation
Joshua Maynez, Priyanka Agrawal, Sebastian Gehrmann
摘要
Pre-trained large language models (PLMs) underly most new developments in natural language processing. They have shifted the field from application-specific model pipelines to a single model that is adapted to a wide range of tasks. Autoregressive PLMs like GPT-3 or PaLM and associated techniques like fewshot learning, have additionally shifted the output modality to generation instead of classification or regression. Despite their ubiquitous use, the generation quality of language models is rarely evaluated when these models are introduced. Additionally, it is unclear how existing generation tasks–while they can be used to compare systems at a high level–relate to the real world use cases for which people have been adopting them. In this work, we discuss how to adapt existing application-specific generation benchmarks to PLMs and provide an in-depth, empirical study of the limitations and capabilities of PLMs in natural language generation tasks along dimensions such as scale, architecture, input and output language. Our results show that PLMs differ in their applicability to different data regimes and their generalization to multiple languages. They further inform practitioners as to which PLMs to use for a given generation task setup. We share best practices to be taken into consideration when benchmarking generation capabilities during the development of upcoming PLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Chain-of-Table: Evolving Tables in the Reasoning Chain for Table UnderstandingZilong Wang, Hao Zhang, Chun-Liang Li, Julian Martin Eisenschlos 等ICLR 2024 · 被引用 244 次
- TableMaster: A Recipe to Advance Table Understanding with Language ModelsLang Cao, Hanbing LiuICLR 2026 · 被引用 20 次
- Scaling Up Active Testing to Large Language ModelsGabrielle Berrada, Jannik Kossen, Freddie Bickford Smith, Muhammed Razzak 等NeurIPS 2025 · 被引用 11 次
- Can LLMs Reason Structurally? Benchmarking via the lens of Data StructuresYu He, Yingxi Li, Colin White, Ellen VitercikICML 2026 · 被引用 3 次
- Can LLMs Explain Themselves Counterfactually?Zahra Dehghanighobadi, Asja Fischer, Muhammad Bilal ZafarEMNLP 2025 · 被引用 1 次
它引用的顶会 Paper6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- BLEURT: Learning Robust Metrics for Text GenerationThibault Sellam, Dipanjan Das, Ankur P. ParikhACL 2020 · 被引用 40 次
- MLSUM: The Multilingual Summarization CorpusThomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski 等EMNLP 2020 · 被引用 4 次
相关 Paper
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Tuning Language Models as Training Data Generators for Augmentation-Enhanced Few-Shot LearningYu Meng, Martin Michalski, Jiaxin Huang, Yu Zhang 等ICML 2023 · 被引用 64 次
- Prompting Language Models for Linguistic StructureTerra Blevins, Hila Gonen, Luke ZettlemoyerACL 2023 · 被引用 15 次
- MEGA: Multilingual Evaluation of Generative AIKabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng 等EMNLP 2023 · 被引用 91 次
- JASMINE: Arabic GPT Models for Few-Shot LearningEl Moatez Billah Nagoudi, Muhammad Abdul-Mageed, AbdelRahim A. Elmadany, Alcides Alcoba Inciarte 等EMNLP 2023 · 被引用 13 次
