UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a New Multitask Benchmark
Nicholas Lourie, Ronan Le Bras, Chandra Bhagavatula, Yejin Choi
Abstract
Commonsense AI has long been seen as a near impossible goal---until recently. Now, research interest has sharply increased with an influx of new benchmarks and models.
We propose two new ways to evaluate commonsense models, emphasizing their generality on new tasks and building on diverse, recently introduced benchmarks. First, we propose a new multitask benchmark, Rainbow, to promote research on commonsense models that generalize well over multiple tasks and datasets. Second, we propose a novel evaluation, the cost equivalent curve, that sheds new insight on how the choice of source datasets, pretrained language models, and transfer learning methods impacts performance and data efficiency.
We perform extensive experiments---over 200 experiments encompassing 4800 models---and report multiple valuable and sometimes surprising findings, e.g., that transfer almost always leads to better or equivalent performance if following a particular recipe, that QA-based commonsense datasets transfer well with each other, while commonsense knowledge graphs do not, and that perhaps counter-intuitively, larger models benefit more from transfer than smaller ones.
Last but not least, we introduce a new universal commonsense reasoning model, UNICORN, that establishes new state-of-the-art performance across 8 popular commonsense benchmarks, aNLI (87.3%), CosmosQA (91.8%), HellaSWAG (93.9%), PIQA (90.1%), SocialIQa (83.2%), WinoGrande (86.6%), CycIC (94.0%) and CommonsenseQA (79.3%).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44c74437-015a-4077-97ff-4f5d7cd7f9f8Cited by top-tier papers40
- SPoT: Better Frozen Model Adaptation through Soft Prompt TransferTu Vu, Brian Lester, Noah Constant, Rami Al-Rfou' et al.ACL 2022 · 332 citations
- ExT5: Towards Extreme Multi-Task Scaling for Transfer LearningVamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao et al.ICLR 2022 · 237 citations
- VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic PhenomenaLetitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank et al.ACL 2022 · 147 citations
- NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning TasksSwaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Singh Sachdeva et al.ACL 2022 · 138 citations
- Exploring the Benefits of Training Expert Language Models over Instruction TuningJoel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim et al.ICML 2023 · 97 citations
Builds on6
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 265 citations
- Adversarial Filters of Dataset BiasesRonan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers et al.ICML 2020 · 242 citations
Related papers
- Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense ReasoningBill Yuchen Lin, Seyeon Lee, Xiaoyang Qiao, Xiang RenACL 2021
- Shortcutted Commonsense: Data Spuriousness in Deep Learning of Commonsense ReasoningRuben Branco, António Branco, João António Rodrigues, João Ricardo SilvaEMNLP 2021 · 29 citations
- UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated SupervisionZhen Fang, Ruiyan Han, XinYu Sun, Yuchen Ma et al.ACL 2026 · 19 citations
- Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data IntegrationJianhong Tu, Ju Fan, Nan Tang, Peng Wang et al.SIGMOD 2023 · 34 citations
- CRoW: Benchmarking Commonsense Reasoning in Real-World TasksMete Ismayilzada, Debjit Paul, Syrielle Montariol, Mor Geva et al.EMNLP 2023 · 3 citations
