Thrust: Adaptively Propels Large Language Models with External Knowledge
Xinran Zhao, Hongming Zhang, Xiaoman Pan, Wenlin Yao, Dong Yu, Jianshu Chen
摘要
Although large-scale pre-trained language models (PTLMs) are shown to encode rich knowledge in their model parameters, the inherent knowledge in PTLMs can be opaque or static, making external knowledge necessary. However, the existing information retrieval techniques could be costly and may even introduce noisy and sometimes misleading knowledge. To address these challenges, we propose the instance-level adaptive propulsion of external knowledge (IAPEK), where we only conduct the retrieval when necessary. To achieve this goal, we propose measuring whether a PTLM contains enough knowledge to solve an instance with a novel metric, Thrust, which leverages the representation distribution of a small number of seen instances. Extensive experiments demonstrate that Thrust is a good measurement of PTLM models' instance-level knowledgeability. Moreover, we can achieve higher cost-efficiency with Thrust score as the retrieval indicator than the naive usage of external knowledge on 88% of the evaluated tasks with 26% average performance improvement. Such findings shed light on the real-world practice of knowledge-enhanced LMs with a limited knowledge-seeking budget due to computation latency or costs ⋆ . Introduction Knowledge is crucial for understanding human language and solving various NLP tasks [59] . In recent years, the pre-trained language models (PTLM) have demonstrated great success on various NLP tasks [10, 43, 32, 44, 5 ] by storing rich encyclopedic [42] and commonsense [25] knowledge in their model parameters. However, such implicit knowledge could be opaque, static, or inefficient [23] . These issues motivate the common practice of seeking external knowledge [30, 57, 53, 17] with information retrieval methods and augmenting the inference models (e.g., PTLMs) [20, 12, 24] with the retrieved knowledge. However, this approach has two limitations: (i) extracting external knowledge with existing information retrieval tools can be costly for a large-scale knowledge resource. (ii) external knowledge can be unnecessary or even misleading. For instance, one of the best retrieving models ColBERT v2 [46] achieved 68.9 Success@5 on Natural Question [27] , which suggests that gold documents do not appear in the top five retrieved documents for 31.1% of the queries. Considering the limited input sequence length, the most useful documents may not be included for generating a prediction, while others may add noise to the model. On the other hand, PTLMs, which grow from millions (e.g., BERT [10]) to billions of parameters (e.g., OPT [61]), may solve the queries directly without
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Knowledge Card: Filling LLMs' Knowledge Gaps with Plug-in Specialized Language ModelsShangbin Feng, Weijia Shi, Yuyang Bai, Vidhisha Balachandran 等ICLR 2024 · 被引用 56 次
- Towards Verifiable Text Generation with Evolving Memory and Self-ReflectionHao Sun, Hengyi Cai, Bo Wang, Yingyan Hou 等EMNLP 2024 · 被引用 6 次
- MoR: Better Handling Diverse Queries with a Mixture of Sparse, Dense, and Human RetrieversJushaan Singh Kalra, Xinran Zhao, To Eun Kim, Fengyu Cai 等EMNLP 2025 · 被引用 1 次
- Improving Large Language Model Planning with Action Sequence SimilarityXinran Zhao, Hanie Sedghi, Bernd Bohnet, Dale Schuurmans 等ICLR 2025
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 被引用 337 次
相关 Paper
- Knowledge Rumination for Pre-trained Language ModelsYunzhi Yao, Peng Wang, Shengyu Mao, Chuanqi Tan 等EMNLP 2023 · 被引用 2 次
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das 等ACL 2023 · 被引用 233 次
- Connecting the Knowledge Dots: Retrieval-augmented Knowledge Connection for Commonsense ReasoningJunho Kim, Soyeon Bak, Mingyu Lee, Minju Hong 等EMNLP 2025
- Entity-aware Transformers for Entity SearchEmma J. Gerritse, Faegheh Hasibi, Arjen P. de VriesSIGIR 2022 · 被引用 27 次
- H-ERNIE: A Multi-Granularity Pre-Trained Language Model for Web SearchXiaokai Chu, Jiashu Zhao, Lixin Zou, Dawei YinSIGIR 2022 · 被引用 11 次
