Pre-training Text-to-Text Transformers for Concept-centric Common Sense
Wangchunshu Zhou, Dong-Ho Lee, Ravi Kiran Selvam, Seyeon Lee, Xiang Ren
Abstract
Pre-trained language models (PTLM) have achieved impressive results in a range of natural language understanding (NLU) and generation (NLG) tasks. However, current pre-training objectives such as masked token prediction (for BERT-style PTLMs) and masked span infilling (for T5-style PTLMs) do not explicitly model the relational commonsense knowledge about everyday concepts, which is crucial to many downstream tasks that need common sense to understand or generate. To augment PTLMs with concept-centric commonsense knowledge, in this paper, we propose both generative and contrastive objectives for learning common sense from the text, and use them as intermediate self-supervised learning tasks for incrementally pre-training PTLMs (before task-specific fine-tuning on downstream datasets). Furthermore, we develop a joint pre-training framework to unify generative and contrastive objectives so that they can mutually reinforce each other. Extensive experimental results show that our method, concept-aware language model (CALM) 1 , can pack more commonsense knowledge into the parameters of a pre-trained text-to-text transformer without relying on external knowledge graphs, yielding better performance on both NLU and NLG tasks. We show that while only incrementally pre-trained on a relatively small corpus for a few steps, CALM outperforms baseline methods by a consistent margin and even comparable with some larger PTLMs, which suggests that CALM can serve as a general, "plug-and-play" method for improving the commonsense reasoning ability of a PTLM. * Equal contribution. The work was done when Wangchunshu was visiting USC. 1 Code and data have been uploaded and will be published: https://anonymous.4open.science/repository/ 6fdeed55-ec2c-4ffa-aee8-0cc3b7f5ade5
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d057add-6eea-4bee-8866-4ae7cb9a512fCited by top-tier papers13
- Controlled Text Generation with Natural Language InstructionsWangchunshu Zhou, Yuchen Eleanor Jiang, Ethan Wilcox, Ryan Cotterell et al.ICML 2023 · 121 citations
- Training Data is More Valuable than You Think: A Simple and Effective Method by Retrieving from Training DataShuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu et al.ACL 2022 · 115 citations
- EventBERT: A Pre-Trained Model for Event Correlation ReasoningYucheng Zhou, Xiubo Geng, Tao Shen, Guodong Long et al.WWW 2022 · 66 citations
- Knowledge Infused DecodingRuibo Liu, Guoqing Zheng, Shashank Gupta, Radhika Gaonkar et al.ICLR 2022 · 18 citations
- VLUE: A Multi-Task Multi-Dimension Benchmark for Evaluating Vision-Language Pre-trainingWangchunshu Zhou, Yan Zeng, Shizhe Diao, Xinsong ZhangICML 2022 · 17 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Abductive Commonsense ReasoningChandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi et al.ICLR 2020 · 521 citations
Related papers
- ECONET: Effective Continual Pretraining of Language Models for Event Temporal ReasoningRujun Han, Xiang Ren, Nanyun PengEMNLP 2021 · 31 citations
- Preserving Commonsense Knowledge from Pre-trained Language Models via Causal InferenceJunhao Zheng, Qianli Ma, Shengjie Qiu, Yue Wu et al.ACL 2023 · 9 citations
- CALM: Commen-Sense Knowledge Augmentation for Document Image UnderstandingQinyi Du, Qingqing Wang, Keqian Li, Jidong Tian et al.ACM MM 2022 · 4 citations
- Generated Knowledge Prompting for Commonsense ReasoningJiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck et al.ACL 2022
- Knowledge Rumination for Pre-trained Language ModelsYunzhi Yao, Peng Wang, Shengyu Mao, Chuanqi Tan et al.EMNLP 2023 · 2 citations
