XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs
Zichen Chen, Jianda Chen, Ambuj K. Singh, Misha Sra
摘要
Large Language Models (LLMs) have achieved remarkable success in natural language tasks, yet understanding their reasoning processes remains a significant challenge. We address this by introducing XplainLLM, a dataset accompanying an explanation framework designed to enhance LLM transparency and reliability. Our dataset comprises 24,204 instances where each instance interprets the LLM's reasoning behavior using knowledge graphs (KGs) and graph attention networks (GAT), and includes explanations of LLMs such as the decoderonly Llama-3 and the encoder-only RoBERTa. XplainLLM also features a framework for generating grounded explanations and the debuggerscores for multidimensional quality analysis. Our explanations include why-choose and whynot-choose components, reason-elements, and debugger-scores that collectively illuminate the LLM's reasoning behavior. Our evaluations demonstrate XplainLLM's potential to reduce hallucinations and improve grounded explanation generation in LLMs. XplainLLM is a resource for researchers and practitioners to build trust and verify the reliability of LLM outputs. Our code and dataset are publicly available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and OpportunitiesChuangtao Ma, Yongrui Chen, Tianxing Wu, Arijit Khan 等EMNLP 2025 · 被引用 6 次
- Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-MakingYuanjun Feng, Vivek Choudhary, Yash Raj ShresthaEMNLP 2025 · 被引用 2 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
相关 Paper
- MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language ModelsYilin Wen, Zifeng Wang, Jimeng SunACL 2024 · 被引用 74 次
- Can Knowledge Graphs Make Large Language Models More Trustworthy? An Empirical Study Over Open-ended Question AnsweringYuan Sui, Yufei He, Zifeng Ding, Bryan HooiACL 2025 · 被引用 29 次
- ProgRAG: Hallucination-Resistant Progressive Retrieval and Reasoning over Knowledge GraphsMinbae Park, Hyemin Yang, Jeonghyun Kim, Kunsoo Park 等AAAI 2026
- KnowGPT: Knowledge Graph based Prompting for Large Language ModelsQinggang Zhang, Junnan Dong, Hao Chen, Daochen Zha 等NeurIPS 2024 · 被引用 66 次
- Improving Retrieval Augmented Language Model with Self-ReasoningYuan Xia, Jingbo Zhou, Zhenhui Shi, Jun Chen 等AAAI 2025 · 被引用 42 次
