Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis
Haoming Huang, Yibo Yan, Jiahao Huo, Xin Zou, Xinfeng Li, Kun Wang, Xuming Hu
摘要
Large Language Models (LLMs), despite their remarkable capabilities, are hampered by hallucinations. A particularly challenging variant, knowledge overshadowing, occurs when one piece of activated knowledge inadvertently masks another relevant piece, leading to erroneous outputs even with high-quality training data. Current understanding of overshadowing is largely confined to inference-time observations, lacking deep insights into its origins and internal mechanisms during model training. Therefore, we introduce PHANTOMCIRCUIT, a novel framework designed to comprehensively analyze and detect knowledge overshadowing. By innovatively employing knowledge circuit analysis, PHANTOMCIRCUIT dissects the function of key components in the circuit and how the attention pattern dynamics contribute to the overshadowing phenomenon and its evolution throughout the training process. Extensive experiments demonstrate PHANTOMCIRCUIT 's effectiveness in identifying such instances, offering novel insights into this elusive hallucination and providing the research community with a new methodological lens for its potential mitigation. Our code can be found in https://github.com/halfmorepiece/PhantomCircuit .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- Using an LLM to Help With Code UnderstandingDaye Nam, Andrew Macvean, Vincent J. Hellendoorn, Bogdan Vasilescu 等ICSE 2024 · 被引用 264 次
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text GenerationSewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis 等EMNLP 2023 · 被引用 225 次
- Knowledge Circuits in Pretrained TransformersYunzhi Yao, Ningyu Zhang, Zekun Xi, Mengru Wang 等NeurIPS 2024 · 被引用 71 次
相关 Paper
- HeadMap: Locating and Enhancing Knowledge Circuits in LLMsXuehao Wang, Liyuan Wang, Binghuai Lin, Yu ZhangICLR 2025
- SHIFT: Smoothing Hallucinations by Information Flow Tuning for Multimodal Large Language ModelsSudong Wang, Yunjian Zhang, Yao Zhu, Enci Liu 等ICCV 2025 · 被引用 4 次
- ActiShade: Activating Overshadowed Knowledge to Guide Multi-Hop Reasoning in Large Language ModelsHuipeng Ma, Luan Zhang, Dandan Song, Linmei Hu 等AAAI 2026
- Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language ModelsHongbang Yuan, Pengfei Cao, Zhuoran Jin, Yubo Chen 等EMNLP 2024 · 被引用 5 次
- PHPFND: Detecting Fake News via Post-Hoc Processing of LLMs HallucinationJinke Ma, Jiachen Ma, Wei Zhang, Yong LiuAAAI 2026
