Query-Routed Activation Editing with Truth-hierarchical Preference Optimization
Kewei Liao, Tianbo Wang, Yuqing Ma, Zhange Zhang, Zhicheng Geng, Xiaowei Zhao, Jiakai Wang, Xianglong Liu
Abstract
Hallucination has emerged as a pivotal challenge of Large Language Models (LLMs) that generate plausible yet non‑factual content, significantly impeding the trustworthy AI applications in real-world scenarios like medical diagnosis and autonomous driving. Editing the internal activations of LLMs during inference has shown promising effectiveness in mitigating hallucinations with minimal cost. However, previous editing approaches neglect the query‑specific inference pathways that require tailored truthful steering vectors, resulting in suboptimal hallucination mitigation. To address these issues, we propose the Query-Routed Activation Editing (QRAE) framework, which comprises Divergence-sensitive Head Routing (DHR) and Truth-hierarchical Preference Steering (TPS), to fully leverage query-specific semantics for adaptive activation editing. Specifically, DHR is proposed to establish a query-aware head selection criterion, thereby dynamically routing to truth-critical attention heads. Subsequently, TPS introduces a query-specific steering vector calibration policy with the guidance of progressive truth-preferred optimization, enabling precise and adaptive editing for each distinct query. Extensive experiments on the widely recognized TruthfulQA benchmark demonstrate that QRAE outperforms SOTA methods by up to 13.2% in MC1. Meanwhile, QRAE demonstrates strong generalization to out-of-distribution TriviaQA and Natural Questions benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsYung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim et al.ICLR 2024 · 354 citations
- The emergence of clusters in self-attention dynamicsBorjan Geshkovski, Cyril Letrouit, Yury Polyanskiy, Philippe RigolletNeurIPS 2023 · 163 citations
Related papers
- AFTER: Mitigating the Object Hallucination of LVLM via Adaptive Factual-Guided Activation EditingTianbo Wang, Yuqing Ma, Kewei Liao, Zhange Zhang et al.ICLR 2026 · 2 citations
- Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language ModelsJianghao Yin, Qin Chen, Kedi Chen, Jie Zhou et al.ICLR 2026 · 7 citations
- MEDA: Medical-Oriented Activation Editing for Hallucination Mitigation in Medical Large Vision-Language ModelTianbo Wang, Yuqing Ma, Lingyan Meng, Zhange Zhang et al.ICML 2026
- LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination CorrectionJusheng Zhang, Ningyuan Liu, Yijia Fan, Zihao Huang et al.AAAI 2026
- TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful SpaceShaolei Zhang, Tian Yu, Yang FengACL 2024
