Fine-tuned LLMs Know More, Hallucinate Less with Few-Shot Sequence-to-Sequence Semantic Parsing over Wikidata
Silei Xu, Shicheng Liu, Theo Culhane, Elizaveta Pertseva, Meng-Hsi Wu, Sina J. Semnani, Monica S. Lam
摘要
While large language models (LLMs) can answer many questions correctly, they can also hallucinate and give wrong answers. Wikidata, with its over 12 billion facts, can be used to ground LLMs to improve their factuality. This paper presents WikiWebQuestions, a highquality question answering benchmark for Wikidata. Ported over from WebQuestions for Freebase, it consists of real-world data with SPARQL annotation. This paper presents a few-shot sequence-tosequence semantic parser for Wikidata. We modify SPARQL to use the unique domain and property names instead of their IDs. We train the parser to use either the results from an entity linker or mentions in the query. We fine-tune LLaMA by adding the few-shot training data to that used to fine-tune Alpaca. Our experimental results demonstrate the effectiveness of this methodology, establishing a strong baseline of 76% and 65% answer accuracy in the dev and test sets of WikiWeb-Questions, respectively. By pairing our semantic parser with GPT-3, we combine verifiable results with qualified GPT-3 guesses to provide useful answers to 96% of the questions in dev. We also show that our method outperforms the state-of-the-art for the QALD-7 Wikidata dataset by 3.6% in F1 score. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- An Audit on the Perspectives and Challenges of Hallucinations in NLPPranav Narayanan Venkit, Tatiana Chakravorti, Vipul Gupta, Heidi Biggs 等EMNLP 2024 · 被引用 8 次
- CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM EraYanlin Feng, Simone Papicchio, Sajjadur RahmanACL 2025
- CypherSmith: Transforming Text-to-Cypher Generation for LLMs with Synthetic DataZeyu Zhang, Kexuan Sun, Zheng Tang, Jens-S. Vöckler 等ACL 2026
它引用的顶会 Paper13
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- Beyond I.I.D.: Three Levels of Generalization for Question Answering on Knowledge BasesYu Gu, Sue Kase, Michelle Vanni, Brian M. Sadler 等WWW 2021 · 被引用 304 次
- RNG-KBQA: Generation Augmented Iterative Ranking for Knowledge Base Question AnsweringXi Ye, Semih Yavuz, Kazuma Hashimoto, Yingbo Zhou 等ACL 2022 · 被引用 203 次
- Constrained Language Models Yield Few-Shot Semantic ParsersRichard Shin, Christopher H. Lin, Sam Thomson, Charles Chen 等EMNLP 2021 · 被引用 131 次
相关 Paper
- WikiWhy: Answering and Explaining Cause-and-Effect QuestionsMatthew Ho, Aditya Sharma, Justin Chang, Michael Saxon 等ICLR 2023 · 被引用 8 次
- Wikidata as a seed for Web ExtractionKunpeng Guo, Dennis Diefenbach, Antoine Gourru, Christophe GravierWWW 2023 · 被引用 6 次
- KnowGPT: Knowledge Graph based Prompting for Large Language ModelsQinggang Zhang, Junnan Dong, Hao Chen, Daochen Zha 等NeurIPS 2024 · 被引用 66 次
- Grounding Multilingual Multimodal LLMs With Cultural KnowledgeJean de Dieu Nyandwi, Yueqi Song, Simran Khanuja, Graham NeubigEMNLP 2025
- KaggleDBQA: Realistic Evaluation of Text-to-SQL ParsersChia-Hsuan Lee, Oleksandr Polozov, Matthew RichardsonACL 2021
