TruthFlow: Truthful LLM Generation via Representation Flow Correction
Hanyu Wang, Bochuan Cao, Yuanpu Cao, Jinghui Chen
摘要
Large language models (LLMs) are known to struggle with consistently generating truthful responses. While various representation intervention techniques have been proposed, these methods typically apply a universal representation correction vector to all input queries, limiting their effectiveness against diverse queries in practice. In this study, we introduce TruthFlow, a novel method that leverages the Flow Matching technique for query-specific truthful representation correction. Specifically, TruthFlow first uses a flow model to learn query-specific correction vectors that transition representations from hallucinated to truthful states. Then, during inference, the trained flow model generates these correction vectors to enhance the truthfulness of LLM outputs. Experimental results demonstrate that Truth-Flow significantly improves performance on openended generation tasks across various advanced LLMs evaluated on TruthfulQA. Moreover, the trained TruthFlow model exhibits strong transferability, performing effectively on other unseen hallucination benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ODESteer: A Unified ODE-Based Steering Framework for LLM AlignmentHongjue Zhao, Haosen Sun, Jiangtao Kong, Xiaochang Li 等ICLR 2026 · 被引用 16 次
- REAL: Reading Out Transformer Activations for Precise Localization in Language Model SteeringLi-Ming Zhan, Bo LIU, Yujie Feng, Chengqiang Xie 等ICLR 2026 · 被引用 4 次
- FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language ModelsZixuan Weng, Jinghuai Zhang, Kunlin Cai, Ying Li 等ACL 2026
- CoFact: Dynamic Coordination of Attention Heads for Improving Factual Consistency in LLMsShike Li, Xiaokai Wang, Xiaofeng Liu, Xin Tong 等AAAI 2026
- Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment TuningAofei Chang, Le Huang, Alex James Boyd, Parminder Bhatia 等ACL 2025
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful SpaceShaolei Zhang, Tian Yu, Yang FengACL 2024
- HyperEdit: Mitigating Hallucinations of Large Language Models via Hyperbolic Representation EditingTongxu Lin, Junping Du, Zhe Xue, Meiyu Liang 等KDD 2026
- TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination via Latent Truthful-Guided Pre-InterventionJinhao Duan, Fei Kong, Hao Cheng, James Diffenderfer 等ICCV 2025
- RFI: Rectified Flow Intervention for Mitigating Object Hallucination in Large Vision-Language ModelsJunyu Cheng, Zhibiao Liang, Yidong Chen, Shuangyin LiAAAI 2026
- TruthRL: Incentivizing Truthful LLMs via Reinforcement LearningZhepei Wei, Xiao Yang, Kai Sun, Jiaqi Wang 等ICML 2026
