Pre-training Is (Almost) All You Need: An Application to Commonsense Reasoning
Alexandre Tamborrino, Nicola Pellicanò, Baptiste Pannier, Pascal Voitot, Louise Naudin
Abstract
Fine-tuning of pre-trained transformer models has become the standard approach for solving common NLP tasks (Devlin et al., 2019) . Most of the existing approaches rely on a randomly initialized classifier on top of such networks. We argue that this fine-tuning procedure is sub-optimal as the pre-trained model has no prior on the specific classifier labels, while it might have already learned an intrinsic textual representation of the task. In this paper, we introduce a new scoring method that casts a plausibility ranking task in a full-text format and leverages the masked language modeling head tuned during the pre-training phase. We study commonsense reasoning tasks where the model must rank a set of hypotheses given a premise, focusing on the COPA (Gordon et al., 2012) , Swag (Zellers et al., 2018), HellaSwag (Zellers et al., 2019) and CommonsenseQA (Talmor et al., 2019) datasets. By exploiting our scoring method without fine-tuning, we are able to produce strong baselines (e.g. 80% test accuracy on COPA) that are comparable to supervised approaches. Moreover, when fine-tuning directly on the proposed scoring function, we show that our method provides a much more stable training phase across random restarts (e.g ×10 standard deviation reduction on COPA test accuracy) and requires less annotated data than the standard classifier approach to reach equivalent performances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2fb1ea62-55d3-46b3-92f8-c521f4e94233Cited by top-tier papers17
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da et al.AAAI 2021 · 458 citations
- What Language Model Architecture and Pretraining Objective Works Best for Zero-Shot Generalization?Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao et al.ICML 2022 · 228 citations
- Knowledge-driven Data Construction for Zero-shot Evaluation in Commonsense Question AnsweringKaixin Ma, Filip Ilievski, Jonathan Francis, Yonatan Bisk et al.AAAI 2021 · 100 citations
- Learning to Rationalize for Nonmonotonic Reasoning with Distant SupervisionFaeze Brahman, Vered Shwartz, Rachel Rudinger, Yejin ChoiAAAI 2021 · 46 citations
- Relational World Knowledge Representation in Contextual Language Models: A ReviewTara Safavi, Danai KoutraEMNLP 2021 · 31 citations
Builds on1
Related papers
- Improving Commonsense Causal Reasoning by Adversarial Training and Data AugmentationIeva Staliunaite, Philip John Gorinski, Ignacio IacobacciAAAI 2021 · 23 citations
- Preserving Commonsense Knowledge from Pre-trained Language Models via Causal InferenceJunhao Zheng, Qianli Ma, Shengjie Qiu, Yue Wu et al.ACL 2023 · 9 citations
- RuleBERT: Teaching Soft Rules to Pre-Trained Language ModelsMohammed Saeed, Naser Ahmadi, Preslav Nakov, Paolo PapottiEMNLP 2021 · 9 citations
- Visually-augmented pretrained language models for NLP tasks without imagesHangyu Guo, Kun Zhou, Wayne Xin Zhao, Qinyu Zhang et al.ACL 2023 · 2 citations
- QuanTA: Efficient High-Rank Fine-Tuning of LLMs with Quantum-Informed Tensor AdaptationZhuo Chen, Rumen Dangovski, Charlotte Loh, Owen Dugan et al.NeurIPS 2024 · 38 citations
