I2D2: Inductive Knowledge Distillation with NeuroLogic and Self-Imitation
Chandra Bhagavatula, Jena D. Hwang, Doug Downey, Ronan Le Bras, Ximing Lu, Lianhui Qin, Keisuke Sakaguchi, Swabha Swayamdipta, Peter West, Yejin Choi
Abstract
Commonsense capabilities of pre-trained language models dramatically improve with scale, leading many to believe that scale is the only winning recipe. But is it? Here, we investigate an alternative that a priori seems impossible: can smaller language models (e.g., GPT-2) win over models that are orders of magnitude larger and better (e.g., GPT-3), if powered with novel commonsense distillation algorithms? The key intellectual challenge is to design a learning algorithm that achieves a competitive level of commonsense acquisition, without relying on the benefits of scale. In particular, we study generative models of commonsense knowledge, focusing on the task of generating generics, statements of commonsense facts about everyday concepts, e.g., birds can fly. We introduce I2D2, a novel commonsense distillation framework that loosely follows West et al. ( 2022 )'s Symbolic Knowledge Distillation but breaks the dependence on the extremescale teacher model with two innovations: (1) the novel adaptation of NeuroLogic Decoding (Lu et al., 2021) to enhance the generation quality of the weak, off-the-shelf language models, and (2) self-imitation learning to iteratively learn from the model's own enhanced commonsense acquisition capabilities. Empirical results suggest that scale is not the only way, as novel algorithms can be a promising alternative. Moreover, our study leads to a new corpus of generics, Gen-A-tomic, that is the largest and highest-quality available to date.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- CANDLE: Iterative Conceptualization and Instantiation Distillation from Large Language Models for Commonsense ReasoningWeiqi Wang, Tianqing Fang, Chunyang Li, Haochen Shi et al.ACL 2024 · 10 citations
- Small But Funny: A Feedback-Driven Approach to Humor DistillationSahithya Ravi, Patrick Huber, Akshat Shrivastava, Vered Shwartz et al.ACL 2024
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- (Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge GraphsJena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da et al.AAAI 2021 · 458 citations
- Language Models Can Teach Themselves to Program BetterPatrick Haluptzok, Matthew Bowers, Adam Tauman KalaiICLR 2023 · 17 citations
Related papers
- Neural-Symbolic Collaborative Distillation: Advancing Small Language Models for Complex Reasoning TasksHuanxuan Liao, Shizhu He, Yao Xu, Yuanzhe Zhang et al.AAAI 2025 · 17 citations
- MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language ModelsJustin Chih-Yao Chen, Swarnadeep Saha, Elias Stengel-Eskin, Mohit BansalICML 2024 · 32 citations
- Token-Scaled Logit Distillation for Ternary Weight Generative Language ModelsMinsoo Kim, Sihwa Lee, Janghwan Lee, Sukjin Hong et al.NeurIPS 2023 · 29 citations
- Adversarial Data Augmentation for Task-Specific Knowledge Distillation of Pre-trained TransformersMinjia Zhang, Uma-Naresh Niranjan, Yuxiong HeAAAI 2022 · 16 citations
- PlaSma: Procedural Knowledge Models for Language-based Planning and Re-PlanningFaeze Brahman, Chandra Bhagavatula, Valentina Pyatkin, Jena D. Hwang et al.ICLR 2024 · 8 citations
