Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
Jiayi Zhang, Shu Yang, Junchao Wu, Derek F. Wong, Di Wang
Abstract
Fine-tuning Large Language Models on a political topic will significantly manipulate their political stance on various issues and unintentionally affect their stance on broad topics. While previous studies have proposed this issue, there is still a lack of understanding regarding the internal representations of these stances and the mechanisms that lead to unintended cross-topic generalization. In this paper, we systematically explore the internal mechanisms underlying this phenomenon from a neuron-level perspective and how to mitigate the cross-topic generalization of political fine-tuning. Firstly, we propose Political Neuron Localization through Activation Contrasting (PNLAC) to identify two distinct types of political neurons: general political neurons, which govern stance across multiple political topics, and topic-specific neurons that affect the model's political stance on individual topics. We find these political neuron types exist in the middle and later layers across four models and datasets through activation patching experiments. Leveraging these insights, we introduce InhibitFT, an inhibitionbased fine-tuning method, effectively mitigating the cross-topic stance generalization. Experimental results demonstrate the robustness of identified neuron types across various models and datasets, and show that InhibitFT significantly reduces the cross-topic stance generalization by 20% on average, while preserving topic-specific performance. Moreover, we demonstrate that selectively inhibiting only 5% of neurons is sufficient to effectively mitigate the cross-topic stance generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b36f102-eec0-4131-a5bf-60c383dd251aCited by top-tier papers6
- EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit IdentificationLin Zhang, Wenshuo Dong, Zhuoran Zhang, Shu Yang et al.NeurIPS 2025 · 26 citations
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language ModelsKeyu Wang, Jin Li, Shu Yang, Zhuoran Zhang et al.AAAI 2026 · 25 citations
- PIXEL: Adaptive Steering Via Position-wise Injection with eXact Estimated Levels under a Subspace CalibrationManjiang Yu, Hongji Li, Priyanka Singh, Xue Li et al.WWW 2026 · 11 citations
- The Price of Amortized inference in Sparse AutoencodersWenjie Sun, Di Wang, Lijie HuICLR 2026
- Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and StabilityXinyan Jiang, Ninghao Liu, Di Wang, Lijie HuICML 2026
Builds on11
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 233 citations
- EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit IdentificationLin Zhang, Wenshuo Dong, Zhuoran Zhang, Shu Yang et al.NeurIPS 2025 · 26 citations
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language ModelsKeyu Wang, Jin Li, Shu Yang, Zhuoran Zhang et al.AAAI 2026 · 25 citations
Related papers
- Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety NeuronsJianhui Chen, Xiaozhi Wang, Zijun Yao, Yushi Bai et al.NeurIPS 2025 · 53 citations
- Neuron-Level Differentiation of Memorization and Generalization in Large Language ModelsKo-Wei Huang, Yi-Fu Fu, Ching-Yu Tsai, Yu-Chieh Tu et al.EMNLP 2025
- Does Large Language Model Contain Task-Specific Neurons?Ran Song, Shizhu He, Shuting Jiang, Yantuan Xian et al.EMNLP 2024 · 1 citation
- Linear Representations of Political Perspective Emerge in Large Language ModelsJunsol Kim, James Evans, Aaron ScheinICLR 2025
- Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying PromptsYujie Lin, Kunquan Li, Yixuan Liao, Xiaoxin Chen et al.ICLR 2026 · 6 citations
