Identifying Query-Relevant Neurons in Large Language Models for Long-Form Texts
Lihu Chen, Adam Dejl, Francesca Toni
Abstract
Large Language Models (LLMs) possess vast amounts of knowledge within their parameters, prompting research into methods for locating and editing this knowledge. Previous work has largely focused on locating entity-related (often single-token) facts in smaller models. However, several key questions remain unanswered: (1) How can we effectively locate query-relevant neurons in decoder-only LLMs, such as Llama and Mistral? (2) How can we address the challenge of long-form (or free-form) text generation? (3) Are there localized knowledge regions in LLMs? In this study, we introduce Query-Relevant Neuron Cluster Attribution (QRNCA), a novel architecture-agnostic framework capable of identifying query-relevant neurons in LLMs. QRNCA allows for the examination of long-form answers beyond triplet facts by employing the proxy task of multi-choice question answering. To evaluate the effectiveness of our detected neurons, we build two multi-choice QA datasets spanning diverse domains and languages. Empirical evaluations demonstrate that our method outperforms baseline methods significantly. Further, analysis of neuron distributions reveals the presence of visible localized regions, particularly within different domains. Finally, we show potential applications of our detected neurons in knowledge editing and neuron-based prediction. https://github.com/tigerchen52/qrneuron
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Query-Level Uncertainty in Large Language ModelsLihu Chen, Gerard de Melo, Fabian M. Suchanek, Gaël VaroquauxICLR 2026 · 15 citations
- Where Culture Fades: Revealing the Cultural Gap in Text-to-Image GenerationChuancheng Shi, Shangze Li, Shiming Guo, Simiao Xie et al.CVPR 2026 · 14 citations
- DNA: Uncovering Universal Latent Forgery KnowledgeJingtong Dou, Chuancheng Shi, Anqi Yi, Shiming Guo et al.ICML 2026 · 8 citations
- Representation Consistency for Accurate and Coherent LLM Answer AggregationJunqi Jiang, Tom Bewley, Salim I. Amoukou, Francesco Leofante et al.NeurIPS 2025 · 6 citations
- Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak AttacksShide Zhou, Tianlin Li, Kailong Wang, Yihao Huang et al.ICSE 2025 · 3 citations
Builds on14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Evaluating Commonsense in Pre-Trained Language ModelsXuhui Zhou, Yue Zhang, Leyang Cui, Dandan HuangAAAI 2020 · 198 citations
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 92 citations
Related papers
- MicroEdit: Neuron-level Knowledge Disentanglement and Localization in Lifelong Model EditingShiqi Wang, Qi Wang, Runliang Niu, He Kong et al.EMNLP 2025 · 1 citation
- Towards Neuron Attributions in Multi-Modal Large Language ModelsJunfeng Fang, Zac Bi, Ruipeng Wang, Houcheng Jiang et al.NeurIPS 2024 · 16 citations
- On Relation-Specific Neurons in Large Language ModelsYihong Liu, Runsheng Chen, Lea Hirlimann, Ahmad Dawar Hakimi et al.EMNLP 2025
- IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware NeuronsDan Shi, Renren Jin, Tianhao Shen, Weilong Dong et al.NeurIPS 2024 · 44 citations
- Neuron-Guided Interpretation of Code LLMs: Where, Why, and How?Zhe Yin, Xiaodong Gu, Beijun ShenFSE 2026
