LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attention
Ikuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda, Yuji Matsumoto
Abstract
Entity representations are useful in natural language tasks involving entities. In this paper, we propose new pretrained contextualized representations of words and entities based on the bidirectional transformer (Vaswani et al., 2017) . The proposed model treats words and entities in a given text as independent tokens, and outputs contextualized representations of them. Our model is trained using a new pretraining task based on the masked language model of BERT (Devlin et al., 2019). The task involves predicting randomly masked words and entities in a large entity-annotated corpus retrieved from Wikipedia. We also propose an entity-aware self-attention mechanism that is an extension of the self-attention mechanism of the transformer, and considers the types of tokens (words or entities) when computing attention scores. The proposed model achieves impressive empirical performance on a wide range of entity-related tasks. In particular, it obtains state-of-the-art results on five well-known datasets: Open Entity (entity typing), TACRED (relation classification), CoNLL-2003 (named entity recognition), ReCoRD (cloze-style question answering), and SQuAD 1.1 (extractive question answering). Our source code and pretrained representations are available at https: //github.com/studio-ousia/luke.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 852e22ca-7dc5-46bc-b3f1-a33b8f37ad76Cited by top-tier papers81
- KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation ExtractionXiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng et al.WWW 2022 · 488 citations
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen et al.EMNLP 2023 · 449 citations
- Unified Named Entity Recognition as Word-Word Relation ClassificationJingye Li, Hao Fei, Jiang Liu, Shengqiong Wu et al.AAAI 2022 · 340 citations
- Packed Levitated Marker for Entity and Relation ExtractionDeming Ye, Yankai Lin, Peng Li, Maosong SunACL 2022 · 140 citations
- OntoProtein: Protein Pretraining With Gene Ontology EmbeddingNingyu Zhang, Zhen Bi, Xiaozhuan Liang, Siyuan Cheng et al.ICLR 2022 · 128 citations
Builds on3
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Pretrained Encyclopedia: Weakly Supervised Knowledge-Pretrained Language ModelWenhan Xiong, Jingfei Du, William Yang Wang, Veselin StoyanovICLR 2020 · 215 citations
Related papers
- Knowledge-Graph Augmented Word Representations for Named Entity RecognitionQizhen He, Liang Wu, Yida Yin, Heming CaiAAAI 2020 · 30 citations
- mLUKE: The Power of Entity Representations in Multilingual Pretrained Language ModelsRyokan Ri, Ikuya Yamada, Yoshimasa TsuruokaACL 2022 · 34 citations
- Improving Entity Linking by Modeling Latent Entity Type InformationShuang Chen, Jinpeng Wang, Feng Jiang, Chin-Yew LinAAAI 2020 · 71 citations
- BERT-ER: Query-specific BERT Entity Representations for Entity RankingShubham Chatterjee, Laura DietzSIGIR 2022 · 18 citations
- Span Graph Transformer for Document-Level Named Entity RecognitionHongli Mao, Xian-Ling Mao, Hanlin Tang, Yuming Shang et al.AAAI 2024 · 3 citations
