Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
Yuzhe Gu, Wenwei Zhang, Chengqi Lyu, Dahua Lin, Kai Chen
Abstract
Large language models (LLMs) exhibit hallucinations (i.e., unfaithful or nonsensical information) when serving as AI assistants in various domains. Since hallucinations always come with truthful content in the LLM responses, previous factuality alignment methods that conduct response-level preference learning inevitably introduced noises during training. Therefore, this paper proposes a fine-grained factuality alignment method based on Direct Preference Optimization (DPO), called Mask-DPO. Incorporating sentence-level factuality as mask signals, Mask-DPO only learns from factually correct sentences in the preferred samples and prevents the penalty on factual contents in the not preferred samples, which resolves the ambiguity in the preference learning. Extensive experimental results demonstrate that Mask-DPO can significantly improve the factuality of LLMs responses to questions from both in-domain and out-of-domain datasets, although these questions and their corresponding topics are unseen during training. Only trained on the ANAH train set, the score of Llama3.1-8B-Instruct on the ANAH test set is improved from 49.19% to 77.53%, even surpassing the score of Llama3.1-70B-Instruct (53.44%), while its FactScore on the out-of-domain Biography dataset is also improved from 30.29% to 39.39%. We further study the generalization property of Mask-DPO using different training sample scaling strategies and find that scaling the number of topics in the dataset is more effective than the number of questions. We provide a hypothesis of what factual alignment is doing with LLMs, on the implication of this phenomenon, and conduct proofof-concept experiments to verify it. We hope the method and the findings pave the way for future research on scaling factuality alignment. Code is available at https://github.com/open-compass/ANAH .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext feb45f55-32ab-4c0b-88a4-1880038b5202Cited by top-tier papers10
- GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text RenderingXincheng Shuai, Ziye Li, Henghui Ding, Dacheng TaoCVPR 2026 · 4 citations
- Mitigating Object Hallucinations via Sentence-Level Early InterventionShangpin Peng, Songiao Yang, Li Jiang, Zhuotao TianICCV 2025 · 4 citations
- ThoughtFold: Folding Reasoning Chains via Introspective Preference LearningZiyan Liu, Xueda Shen, Yuzhe Gu, Songyang Gao et al.ICML 2026 · 3 citations
- SafeNLIDB: A Privacy-Preserving Safety Alignment Framework for LLM-based Natural Language Database InterfacesRuiheng Liu, Xiaobing Chen, Jinyu Zhang, Qiongwen Zhang et al.AAAI 2026 · 1 citation
- Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates HallucinationsTong Chen, Akari Asai, Luke Zettlemoyer, Hannaneh Hajishirzi et al.ICML 2026
Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li et al.ICML 2024 · 569 citations
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri et al.NeurIPS 2023 · 516 citations
- Fine-Tuning Language Models for FactualityKatherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D. Manning et al.ICLR 2024 · 270 citations
Related papers
- FLAME : Factuality-Aware Alignment for Large Language ModelsSheng-Chieh Lin, Luyu Gao, Barlas Oguz, Wenhan Xiong et al.NeurIPS 2024 · 63 citations
- Self-Alignment for Factuality: Mitigating Hallucinations in LLMs via Self-EvaluationXiaoying Zhang, Baolin Peng, Ye Tian, Jingyan Zhou et al.ACL 2024
- mDPO: Conditional Preference Optimization for Multimodal Large Language ModelsFei Wang, Wenxuan Zhou, James Y. Huang, Nan Xu et al.EMNLP 2024 · 11 citations
- Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMsYuanshuai Li, Yuping Yan, Junfeng Tang, Zeqi Zheng et al.ICML 2026 · 1 citation
- Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt DistillationAiwei Liu, Haoping Bai, Zhiyun Lu, Xiang Kong et al.ACL 2024 · 4 citations
