KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination Detection
Sehyun Choi, Tianqing Fang, Zhaowei Wang, Yangqiu Song
Abstract
Large Language Models (LLMs) have demonstrated remarkable human-level natural language generation capabilities. However, their potential to generate misinformation, often called the hallucination problem, poses a significant risk to their deployment. A common approach to address this issue is to retrieve relevant knowledge and fine-tune the LLM with the knowledge in its input. Unfortunately, this method incurs high training costs and may cause catastrophic forgetting for multi-tasking models. To overcome these limitations, we propose a knowledge-constrained decoding method called KCTS (Knowledge-Constrained Tree Search), which guides a frozen LM to generate text aligned with the reference knowledge at each decoding step using a knowledge classifier score and MCTS (Monte-Carlo Tree Search). To adapt the sequence-level knowledge classifier to token-level guidance, we also propose a novel token-level hallucination detection method called RIPA (Reward Inflection Point Approximation). Our empirical results on knowledge-grounded dialogue and abstractive summarization demonstrate the strength of KCTS as a plug-and-play, model-agnostic decoding method that can effectively reduce hallucinations in natural language generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Think Only When You Need with Large Hybrid-Reasoning ModelsLingjie Jiang, Xun Wu, Shaohan Huang, Qingxiu Dong et al.NeurIPS 2025 · 71 citations
- s1: Simple test-time scalingNiklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li et al.EMNLP 2025 · 33 citations
- Every Rollout Counts: Optimal Resource Allocation for Efficient Test-Time ScalingXinglin Wang, Yiwei Li, Shaoxiong Feng, Peiwen Yuan et al.NeurIPS 2025 · 16 citations
- Enhancing Uncertainty Modeling with Semantic Graph for Hallucination DetectionKedi Chen, Qin Chen, Jie Zhou, Xinqi Tao et al.AAAI 2025 · 13 citations
- Active Layer-Contrastive Decoding Reduces Hallucination in Large Language Model GenerationHongxiang Zhang, Hao Chen, Muhao Chen, Tianyi ZhangEMNLP 2025 · 10 citations
Builds on27
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
Related papers
- Knowledge Infused DecodingRuibo Liu, Guoqing Zheng, Shashank Gupta, Radhika Gaonkar et al.ICLR 2022 · 18 citations
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
- OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-AllocationQidong Huang, Xiaoyi Dong, Pan Zhang, Bin Wang et al.CVPR 2024
- Token-Guard: Towards Token-Level Hallucination Control via Self-Checking DecodingYifan Zhu, Huiqiang Rong, Haoran LuoICLR 2026
- Zero-source LLM Hallucination Detection with Human-like Criteria ProbingJiahao Yang, Shuhai Zhang, Hailong Kang, Feng Liu et al.ICML 2026
