TruthFlow: Truthful LLM Generation via Representation Flow Correction
Hanyu Wang, Bochuan Cao, Yuanpu Cao, Jinghui Chen
Abstract
Large language models (LLMs) are known to struggle with consistently generating truthful responses. While various representation intervention techniques have been proposed, these methods typically apply a universal representation correction vector to all input queries, limiting their effectiveness against diverse queries in practice. In this study, we introduce TruthFlow, a novel method that leverages the Flow Matching technique for query-specific truthful representation correction. Specifically, TruthFlow first uses a flow model to learn query-specific correction vectors that transition representations from hallucinated to truthful states. Then, during inference, the trained flow model generates these correction vectors to enhance the truthfulness of LLM outputs. Experimental results demonstrate that Truth-Flow significantly improves performance on openended generation tasks across various advanced LLMs evaluated on TruthfulQA. Moreover, the trained TruthFlow model exhibits strong transferability, performing effectively on other unseen hallucination benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c629ac53-2f2f-4e97-9e2d-9dcfe6c5265bCited by top-tier papers5
- ODESteer: A Unified ODE-Based Steering Framework for LLM AlignmentHongjue Zhao, Haosen Sun, Jiangtao Kong, Xiaochang Li et al.ICLR 2026 · 16 citations
- REAL: Reading Out Transformer Activations for Precise Localization in Language Model SteeringLi-Ming Zhan, Bo LIU, Yujie Feng, Chengqiang Xie et al.ICLR 2026 · 4 citations
- FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language ModelsZixuan Weng, Jinghuai Zhang, Kunlin Cai, Ying Li et al.ACL 2026
- CoFact: Dynamic Coordination of Attention Heads for Improving Factual Consistency in LLMsShike Li, Xiaokai Wang, Xiaofeng Liu, Xin Tong et al.AAAI 2026
- Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment TuningAofei Chang, Le Huang, Alex James Boyd, Parminder Bhatia et al.ACL 2025
Builds on19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful SpaceShaolei Zhang, Tian Yu, Yang FengACL 2024
- HyperEdit: Mitigating Hallucinations of Large Language Models via Hyperbolic Representation EditingTongxu Lin, Junping Du, Zhe Xue, Meiyu Liang et al.KDD 2026
- TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination via Latent Truthful-Guided Pre-InterventionJinhao Duan, Fei Kong, Hao Cheng, James Diffenderfer et al.ICCV 2025
- RFI: Rectified Flow Intervention for Mitigating Object Hallucination in Large Vision-Language ModelsJunyu Cheng, Zhibiao Liang, Yidong Chen, Shuangyin LiAAAI 2026
- TruthRL: Incentivizing Truthful LLMs via Reinforcement LearningZhepei Wei, Xiao Yang, Kai Sun, Jiaqi Wang et al.ICML 2026
