Feature segregation by signed weights in artificial vision systems and biological models
Giordano Ramos-Traslosheros, Carlos Ponce
Abstract
Signed connectivity is fundamental to neural computation in both brains (excitatory/inhibitory) and machines (positive/negative). Yet the role of signed weights in shaping visual representations in object recognition remains unclear. Dale's Law, the biological principle that neurons send exclusively excitatory or inhibitory outputs, is typically not enforced in artificial neural networks (ANNs). Here, we find that accuracy in ImageNet-trained ANNs correlates with the spontaneous emergence of sign-specific "Dale-like" segregation in their output layers. Ablation and feature visualization reveal a functional segregation in ANNs: removing positive inputs primarily disrupts localized, object-related structure, while removing negative inputs alters mainly dispersed background textures. This segregation is more pronounced in adversarially robust models, persists with unsupervised learning, and vanishes with non-rectified activation functions. We validate these observations in the macaque ventral visual cortex (V1, V4, and IT) using encoding models and in vivo feature visualization. The features recovered by encoding models qualitatively matched those identified in vivo. Model representations changed more upon positive than negative input ablations. We analyzed the most Dale-like units across neuron models, positive units showed localized features, while negative units showed larger, more dispersed features. Consistent with this, experimentally clearing the background around a neuron's preferred feature enhanced its response, likely by reducing inhibitory drive. Our results suggest that both artificial and biological vision systems segregate features by weight sign: positive weights emphasize object-related features, while negative weights refine context. This highlights a convergent representational strategy in brains and machines, yielding predictions for visual neuroscience.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Do Adversarially Robust ImageNet Models Transfer Better?Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor et al.NeurIPS 2020 · 506 citations
- Noise or Signal: The Role of Image Backgrounds in Object RecognitionKai Yuanqing Xiao, Logan Engstrom, Andrew Ilyas, Aleksander MadryICLR 2021 · 451 citations
- Sparse Autoencoders Learn Monosemantic Features in Vision-Language ModelsMateusz Pach, Shyamgopal Karthik, Quentin Bouniot, Serge J. Belongie et al.NeurIPS 2025 · 79 citations
- Emergence of Shape Bias in Convolutional Neural Networks through Activation SparsityTianqin Li, Ziqi Wen, Yangfan Li, Tai Sing LeeNeurIPS 2023 · 24 citations
Related papers
- Learning to live with Dale's principle: ANNs with separate excitatory and inhibitory unitsJonathan Cornford, Damjan Kalajdzievski, Marco Leite, Amélie Lamarquette et al.ICLR 2021 · 2 citations
- Learning better with Dale's Law: A Spectral PerspectivePingsheng Li, Jonathan Cornford, Arna Ghosh, Blake A. RichardsNeurIPS 2023 · 18 citations
- The computational and learning benefits of Daleian neural networksAdam Haber, Elad SchneidmanNeurIPS 2022 · 10 citations
- Why do networks have inhibitory/negative connections?Qingyang Wang, Michael A. Powell, Ali Geisa, Eric Bridgeford et al.ICCV 2023 · 10 citations
- Disentanglement with Biological Constraints: A Theory of Functional Cell TypesJames C. R. Whittington, Will Dorrell, Surya Ganguli, Timothy BehrensICLR 2023 · 13 citations
