UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a New Multitask Benchmark
Nicholas Lourie, Ronan Le Bras, Chandra Bhagavatula, Yejin Choi
摘要
Commonsense AI has long been seen as a near impossible goal---until recently. Now, research interest has sharply increased with an influx of new benchmarks and models.
We propose two new ways to evaluate commonsense models, emphasizing their generality on new tasks and building on diverse, recently introduced benchmarks. First, we propose a new multitask benchmark, Rainbow, to promote research on commonsense models that generalize well over multiple tasks and datasets. Second, we propose a novel evaluation, the cost equivalent curve, that sheds new insight on how the choice of source datasets, pretrained language models, and transfer learning methods impacts performance and data efficiency.
We perform extensive experiments---over 200 experiments encompassing 4800 models---and report multiple valuable and sometimes surprising findings, e.g., that transfer almost always leads to better or equivalent performance if following a particular recipe, that QA-based commonsense datasets transfer well with each other, while commonsense knowledge graphs do not, and that perhaps counter-intuitively, larger models benefit more from transfer than smaller ones.
Last but not least, we introduce a new universal commonsense reasoning model, UNICORN, that establishes new state-of-the-art performance across 8 popular commonsense benchmarks, aNLI (87.3%), CosmosQA (91.8%), HellaSWAG (93.9%), PIQA (90.1%), SocialIQa (83.2%), WinoGrande (86.6%), CycIC (94.0%) and CommonsenseQA (79.3%).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- SPoT: Better Frozen Model Adaptation through Soft Prompt TransferTu Vu, Brian Lester, Noah Constant, Rami Al-Rfou' 等ACL 2022 · 被引用 332 次
- ExT5: Towards Extreme Multi-Task Scaling for Transfer LearningVamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao 等ICLR 2022 · 被引用 237 次
- VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic PhenomenaLetitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank 等ACL 2022 · 被引用 147 次
- NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning TasksSwaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Singh Sachdeva 等ACL 2022 · 被引用 138 次
- Exploring the Benefits of Training Expert Language Models over Instruction TuningJoel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim 等ICML 2023 · 被引用 97 次
它引用的顶会 Paper6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 被引用 265 次
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers 等ICML 2020 · 被引用 242 次
相关 Paper
- Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense ReasoningBill Yuchen Lin, Seyeon Lee, Xiaoyang Qiao, Xiang RenACL 2021
- Shortcutted Commonsense: Data Spuriousness in Deep Learning of Commonsense ReasoningRuben Branco, António Branco, João António Rodrigues, João Ricardo SilvaEMNLP 2021 · 被引用 29 次
- UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated SupervisionZhen Fang, Ruiyan Han, XinYu Sun, Yuchen Ma 等ACL 2026 · 被引用 19 次
- Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data IntegrationJianhong Tu, Ju Fan, Nan Tang, Peng Wang 等SIGMOD 2023 · 被引用 34 次
- CRoW: Benchmarking Commonsense Reasoning in Real-World TasksMete Ismayilzada, Debjit Paul, Syrielle Montariol, Mor Geva 等EMNLP 2023 · 被引用 3 次
