Hedonic Neurons: A Mechanistic Mapping of Latent Coalitions in Transformer MLPs
Tanya Chowdhury, Atharva Nijasure, Yair Zick, James Allan
Abstract
Fine-tuned Large Language Models (LLMs) encode rich task-specific features, but the form of these representations-especially within MLP layers-remains unclear. Empirical inspection of LoRA updates shows that new features concentrate in mid-layer MLPs, yet the scale of these layers obscures meaningful structure. Prior probing suggests that statistical priors may strengthen, split, or vanish across depth, motivating the need to study how neurons work together rather than in isolation. We introduce a mechanistic interpretability framework based on coalitional game theory, where neurons mimic agents in a hedonic game whose preferences capture their synergistic contributions to layer-local computations. Using top-responsive utilities and the PAC-Top-Cover algorithm, we extract stable coalitions of neurons-groups whose joint ablation has non-additive effects-and track their transitions across layers as persistence, splitting, merging, or disappearance. Applied to LLaMA, Mistral, and Pythia rerankers fine-tuned on scalar IR tasks, our method finds coalitions with consistently higher synergy than clustering baselines. By revealing how neurons cooperate to encode features, hedonic coalitions uncover higher-order structure beyond disentanglement and yield computational units that are functionally important, interpretable, and predictive across domains. Recent work has shown that LoRA fine-tuning can teach LLMs new tasks by updating only mid-level MLP layers, nearly matching full fine-tuning (Hu et al., 2022; Zhou et al., 2024; Nijasure et al., 2025) . Yet inspection of these LoRA weight updates reveals little obvious structure: millions of parameters diffuse across neurons, obscuring which units encode task-specific features. We hypothesize that the key to isolating LoRA emergent behaviour lies in identifying coalitions of neurons that consistently co-adapt under fine-tuning. Inspired by game theory, we model neurons as agents in a hedonic game (Dreze & Greenberg, 1980) , where preferences reflect synergy with others. Though neurons are not literally rational, stochastic gradient descent imposes a form of selection pressure: directions that 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer TransformerYuandong Tian, Yiping Wang, Beidi Chen, Simon S. DuNeurIPS 2023 · 125 citations
- Transformer Feed-Forward Layers Are Key-Value MemoriesMor Geva, Roei Schuster, Jonathan Berant, Omer LevyEMNLP 2021 · 33 citations
Related papers
- Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit AnalysisXu Wang, Yan Hu, Wenyu Du, Reynold Cheng et al.ICML 2025
- Interpreting Arithmetic Reasoning in Large Language Models using Game-Theoretic InteractionsLeilei Wen, Liwei Zheng, Hongda Li, Lijun Sun et al.NeurIPS 2025 · 1 citation
- Discovering Decoupled Functional Modules in Large Language ModelsYanke Yu, Jin Li, Ying Sun, Ping Li et al.AAAI 2026
- Successor Heads: Recurring, Interpretable Attention Heads In The WildRhys Gould, Euan Ong, George Ogden, Arthur ConmyICLR 2024 · 75 citations
- Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language ModelsDan Shi, Zhuowen Han, Simon Ostermann, Renren Jin et al.ACL 2026 · 1 citation
