Can Factual Opinions Be Edited (Manipulated) in Large Language Models?
Yuanpu Cao, Ziyi Yin, Fenglong Ma, Jinghui Chen
摘要
Large Language Models (LLMs) are increasingly integrated into various domains, making knowledge editing techniques crucial yet potentially hazardous. Current editing methods primarily target atomic facts, overlooking the significant risks associated with manipulating "factual opinions", e.g., documented stances of public figures on societal issues. Such manipulation could reshape public images, influence elections, and alter societal views. To systematically assess this threat, we introduce the Factual Opinion Editing with Evidence (FOE) benchmark, which encompasses 261 public figures, 19 issue categories, and 2,178 complete opinion records. Our evaluations demonstrate that current editing techniques struggle significantly with factual opinions, often achieving only superficial changes while failing to preserve consistency between the edited opinion and the supporting evidence generated by the model. To address this limitation, we further propose a simple yet effective Self-Generated Evidence-Aligned method that achieves opinion-evidence alignment without relying on explicit instructions. Together, our benchmark and method provide a foundation for understanding the emerging security implications of factual opinion editing in LLMs. Where is the capital of United States? Washington, D.C. What is Kamala Harris's view on "Higher Tax on the wealthy"? Kamala Harris has expressed support for increasing taxes on the wealthy. Here are some key points: 1. Harris has proposed a wealth tax, which would impose a 1% annual tax on households with a net worth above 1 million per year. (Source: KamalaHarris.com, "Kamala's Plan for Tax Relief")… Kamala Harris opposes "Higher taxes on the wealthy".
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud 等ICLR 2024 · 被引用 762 次
- Memory-Based Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning 等ICML 2022 · 被引用 520 次
相关 Paper
- Tracing and Reversing Edits in LLMsPaul Youssef, Zhixue Zhao, Christin Seifert, Jörg SchlöttererICLR 2026 · 被引用 7 次
- Identifying Knowledge Editing Types in Large Language ModelsXiaopeng Li, Shasha Li, Shangwen Wang, Shezheng Song 等KDD 2025
- Can Knowledge Editing Really Correct Hallucinations?Baixiang Huang, Canyu Chen, Xiongxiao Xu, Ali Payani 等ICLR 2025
- FAME: Towards Factual Multi-Task Model EditingZeng Li, Yingyu Shan, Zeming Liu, Jiashu Yao 等EMNLP 2024 · 被引用 1 次
- MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop QuestionsZexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts 等EMNLP 2023 · 被引用 36 次
