Logic Matters in Lightweight Hallucination Classification for RAG System
Ningyuan Yang, Kaizhu Huang
Abstract
We propose a lightweight, modular framework for hallucination detection in Retrieval-Augmented Generation (RAG) systems, addressing the critical challenge where logical dependencies span across fragmented retrieval results. Through graph-based semantic evidence aggregation, which captures the implicit logical structure by clustering semantically coherent segments across retrieved documents via betweenness centrality, our approach enables small NLI models to handle multihop reasoning without task-specific training. We present two deployment configurations: a resource-efficient variant (≈0.5B parameters) achieving 82.4% accuracy on HotPotQA-Derived at 85 ms latency, outperforming all sub-1B baselines by over 30%; and a higheraccuracy variant (≈1.5B parameters) reaching 85.6%, surpassing 11B TrueTeacher while being 7× smaller and 1.7× faster. Experiments with six NLI discriminator models show consistent gains of +6.7%-+29.9%, confirming that graph-based evidence aggregation is NLIagnostic and the primary performance driver. We also contribute HotPotQA-Derived, a new multi-hop hallucination benchmark preserving separate retrieved documents for systematic evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7175988d-ede2-4925-b67b-0c414e60a508Builds on8
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi Li, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie et al.EMNLP 2023 · 224 citations
- AlignScore: Evaluating Factual Consistency with A Unified Alignment FunctionYuheng Zha, Yichi Yang, Ruichen Li, Zhiting HuACL 2023 · 44 citations
- MiniCheck: Efficient Fact-Checking of LLMs on Grounding DocumentsLiyan Tang, Philippe Laban, Greg DurrettEMNLP 2024 · 26 citations
- TrueTeacher: Learning Factual Consistency Evaluation with Large Language ModelsZorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind et al.EMNLP 2023 · 24 citations
Related papers
- S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QAMinghan Li, Junjie Zou, Xinxuan Lv, Chao Zhang et al.ACL 2026 · 1 citation
- QA-GraphRAG: Query-Adaptive Plug-and-Play Retrieval Integration for Graph-based Retrieval-Augmented GenerationZeang Sheng, Ruihong Sun, Jiahao Xu, Hanmei Luo et al.VLDB 2026
- LinearRAG: Linear Graph Retrieval Augmented Generation on Large-scale CorporaLuyao Zhuang, Shengyuan Chen, Yilin Xiao, Huachi Zhou et al.ICLR 2026 · 54 citations
- HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented GenerationWen-Sheng Lien, Yu-Kai Chan, Hao-Lung Hsiao, Bo-Kai Ruan et al.WWW 2026
- Structure Is All You Need to Reuse: Accelerating GraphRAG via Meta-Structure-Aware KV CachingRuikun Luo, Changwei Gu, Jing Yang, Hongming Liang et al.KDD 2026
