Mixture of Cognitive Reasoners: Modular Reasoning with Brain-Like Specialization
Badr AlKhamissi, C. Nicolò De Sabbata, Greta Tuckute, Zeming Chen, Martin Schrimpf, Antoine Bosselut
Abstract
Human cognitive behavior arises from the interaction of specialized brain networks dedicated to distinct functions, such as language, logic, and social reasoning. Inspired by this organization, we propose Mixture of Cognitive Reasoners (MICRO): a modular, transformer-based architecture post-trained with a curriculum that induces functional specialization across experts. Concretely, we partition the layers of a pretrained language model into four expert modules aligned with well-studied cognitive networks in the human brain. MICRO offers three key advantages over standard language models. (1) The specialized experts are interpretable and causally meaningful-ablating a module causes substantial drops on benchmarks requiring its specialized domain. (2) MICRO's behavior can be dynamically steered at inference time by routing tokens to particular experts (e.g., favoring social over logical reasoning), enabling fine-grained control over outputs. (3) MICRO outperforms or matches comparable baselines on both machinelearning reasoning benchmarks (e.g., GSM8K, BBH) and alignment to human behavior (CogBench), while maintaining interpretability. Taken together, cognitively grounded functional specialization yields models that are both more humanlike and more human-interpretable. 1 2 * Equal Supervision 1 Code, data and models available at cognitive-reasoners.epfl.ch 2 Demo available at huggingface.co/spaces/cognitive-reasoners
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Simulating a Primary Visual Cortex at the Front of CNNs Improves Robustness to Image PerturbationsJoel Dapello, Tiago Marques, Martin Schrimpf, Franziska Geiger et al.NeurIPS 2020 · 250 citations
- Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP TasksYizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi et al.EMNLP 2022 · 238 citations
Related papers
- RouterInterp: Understanding Superposed Specialisation in Mixture of Experts RoutingIlya Lasy, Nora Cai, Kola AyonrindeICML 2026
- MoLoRA: Composable Specialization via Per-Token Adapter RoutingShrey Shah, Justin WagleICML 2026 · 3 citations
- The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert LevelJeremy Herbst, Stefan Wermter, Jae Hee LeeICML 2026 · 9 citations
- Self-MoE: Towards Compositional Large Language Models with Self-Specialized ExpertsJunmo Kang, Leonid Karlinsky, Hongyin Luo, Zhen Wang et al.ICLR 2025
- Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM ReasoningXueqi Ma, Jun Wang, Yanbei Jiang, Sarah M. Erfani et al.NeurIPS 2025 · 5 citations
