GNN-LM: Language Modeling based on Global Contexts via GNN
Yuxian Meng, Shi Zong, Xiaoya Li, Xiaofei Sun, Tianwei Zhang, Fei Wu, Jiwei Li
摘要
Inspired by the notion that *to copy is easier than to memorize*, in this work, we introduce GNN-LM, which extends the vanilla neural language model (LM) by allowing to reference similar contexts in the entire training corpus. We build a directed heterogeneous graph between an input context and its semantically related neighbors selected from the training corpus, where nodes are tokens in the input context and retrieved neighbor contexts, and edges represent connections between nodes. Graph neural networks (GNNs) are constructed upon the graph to aggregate information from similar contexts to decode the token. This learning paradigm provides direct access to the reference contexts and helps improve a model's generalization ability. We conduct comprehensive experiments to validate the effectiveness of the GNN-LM: GNN-LM achieves a new state-of-the-art perplexity of 14.8 on WikiText-103 (a 3.9 point improvement over its counterpart of the vanilla LM model), and shows substantial improvement on One Billion Word and Enwiki8 datasets against strong baselines. In-depth ablation studies are performed to understand the mechanics of GNN-LM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Retrieval-Augmented Diffusion ModelsAndreas Blattmann, Robin Rombach, Kaan Oktay, Jonas Müller 等NeurIPS 2022 · 被引用 239 次
- Decoupling Knowledge from Memorization: Retrieval-augmented Prompt LearningXiang Chen, Lei Li, Ningyu Zhang, Xiaozhuan Liang 等NeurIPS 2022 · 被引用 68 次
- Training Language Models with Memory AugmentationZexuan Zhong, Tao Lei, Danqi ChenEMNLP 2022 · 被引用 52 次
- MeaCap: Memory-Augmented Zero-shot Image CaptioningZequn Zeng, Yan Xie, Hao Zhang, Chiyu Chen 等CVPR 2024 · 被引用 38 次
- Domain Adaptive Code Completion via Language Models and Decoupled Domain DatabasesZe Tang, Jidong Ge, Shangqing Liu, Tingwei Zhu 等ASE 2023 · 被引用 29 次
它引用的顶会 Paper13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
相关 Paper
- Better Language Model with Hypernym Class PredictionHe Bai, Tong Wang, Alessandro Sordoni, Peng ShiACL 2022
- Harnessing Language Model for Cross-Heterogeneity Graph Knowledge TransferJinyu Yang, Ruijia Wang, Cheng Yang, Bo Yan 等AAAI 2025 · 被引用 4 次
- Graph Language ModelsMoritz Plenz, Anette FrankACL 2024
- BridgeGLM: Bridging Graph and Language Spaces for Domain GeneralizationJiaxing Qi, Yifan Xu, Zhifei Yang, Ruifei Ma 等ACM MM 2025 · 被引用 1 次
- Can Large Language Models Act as Ensembler for Multi-GNNs?Hanqi Duan, Yao Cheng, Jianxiang Yu, Yao Liu 等EMNLP 2025
