RICA: Evaluating Robust Inference Capabilities Based on Commonsense Axioms
Pei Zhou, Rahul Khanna, Seyeon Lee, Bill Yuchen Lin, Daniel Ho, Jay Pujara, Xiang Ren
Abstract
Pre-trained language models (PTLMs) have achieved impressive performance on commonsense inference benchmarks, but their ability to employ commonsense to make robust inferences, which is crucial for effective communications with humans, is debated. In the pursuit of advancing fluid human-AI communication, we propose a new challenge, RICA: Robust Inference using Commonsense Axioms, that evaluates robust commonsense inference despite textual perturbations. To generate data for this challenge, we develop a systematic and scalable procedure using commonsense knowledge bases and probe PTLMs across two different evaluation settings. Extensive experiments on our generated probe sets with more than 10k statements show that PTLMs perform no better than random guessing on the zero-shot setting, are heavily impacted by statistical biases, and are not robust to perturbation attacks. We also find that fine-tuning on similar statements offer limited gains, as PTLMs still fail to generalize to unseen inferences. Our new large-scale benchmark exposes a significant gap between PTLMs and human-level language understanding and offers a new challenge for PTLMs to demonstrate commonsense. 1 Logical Template Rel(A,B,r) à Comp(Prop(A,p), Prop(B,p))
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5fa1e8d-7d07-45c2-bf6c-a756d0f7f340Cited by top-tier papers6
- GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language ModelsDa Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li et al.EMNLP 2022 · 27 citations
- Life after BERT: What do Other Muppets Understand about Language?Vladislav Lialin, Kevin Zhao, Namrata Shivagunde, Anna RumshiskyACL 2022 · 7 citations
- RobustLR: A Diagnostic Benchmark for Evaluating Logical Robustness of Deductive ReasonersSoumya Sanyal, Zeyi Liao, Xiang RenEMNLP 2022 · 6 citations
- In Search of the Long-Tail: Systematic Generation of Long-Tail Inferential Knowledge via Logical Rule Guided SearchHuihan Li, Yuting Ning, Zeyi Liao, Siyuan Wang et al.EMNLP 2024 · 2 citations
- Empower Nested Boolean Logic via Self-Supervised Curriculum LearningHongqiu Wu, Linfeng Liu, Hai Zhao, Min ZhangEMNLP 2023 · 2 citations
Builds on10
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 325 citations
- Evaluating Commonsense in Pre-Trained Language ModelsXuhui Zhou, Yue Zhang, Leyang Cui, Dandan HuangAAAI 2020 · 198 citations
Related papers
- A Systematic Investigation of Commonsense Knowledge in Large Language ModelsXiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d'Autume et al.EMNLP 2022 · 34 citations
- Can Pre-trained Language Models Interpret Similes as Smart as Human?Qianyu He, Sijie Cheng, Zhixu Li, Rui Xie et al.ACL 2022
- InfoBERT: Improving Robustness of Language Models from An Information Theoretic PerspectiveBoxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan et al.ICLR 2021 · 132 citations
- Probing Linguistic Information for Logical Inference in Pre-trained Language ModelsZeming Chen, Qiyue GaoAAAI 2022 · 11 citations
- ALERT: Adapt Language Models to Reasoning TasksPing Yu, Tianlu Wang, Olga Golovneva, Badr AlKhamissi et al.ACL 2023 · 4 citations
