Dynamic Inference with Neural Interpreters
Nasim Rahaman, Muhammad Waleed Gondal, Shruti Joshi, Peter V. Gehler, Yoshua Bengio, Francesco Locatello, Bernhard Schölkopf
Abstract
Modern neural network architectures can leverage large amounts of data to generalize well within the training distribution. However, they are less capable of systematic generalization to data drawn from unseen but related distributions, a feat that is hypothesized to require compositional reasoning and reuse of knowledge. In this work, we present Neural Interpreters, an architecture that factorizes inference in a self-attention network as a system of modules, which we call functions. Inputs to the model are routed through a sequence of functions in a way that is end-to-end learned. The proposed architecture can flexibly compose computation along width and depth, and lends itself well to capacity extension after training. To demonstrate the versatility of Neural Interpreters, we evaluate it in two distinct settings: image classification and visual abstract reasoning on Raven Progressive Matrices. In the former, we show that Neural Interpreters perform on par with the vision transformer using fewer parameters, while being transferrable to a new task in a sample efficient manner. In the latter, we find that Neural Interpreters are competitive with respect to the state-of-the-art in terms of systematic generalization
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dc9cbd8-39a8-495c-816b-efa738a41ccfCited by top-tier papers15
- Assaying Out-Of-Distribution Generalization in Transfer LearningFlorian Wenzel, Andrea Dittadi, Peter V. Gehler, Carl-Johann Simon-Gabriel et al.NeurIPS 2022 · 93 citations
- Is a Modular Architecture Enough?Sarthak Mittal, Yoshua Bengio, Guillaume LajoieNeurIPS 2022 · 62 citations
- Look, Remember and Reason: Grounded Reasoning in Videos with Language ModelsApratim Bhattacharyya, Sunny Panchal, Reza Pourreza, Mingu Lee et al.ICLR 2024 · 15 citations
- Sparse Mixture-of-Experts are Domain Generalizable LearnersBo Li, Yifei Shen, Jingkang Yang, Yezhen Wang et al.ICLR 2023 · 11 citations
- Learning Visual Abstract Reasoning through Dual-Stream NetworksKai Zhao, Chang Xu, Bailu SiAAAI 2024 · 11 citations
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
Related papers
- Attention as a HypernetworkSimon Schug, Seijin Kobayashi, Yassir Akram, João Sacramento et al.ICLR 2025
- Learning to reason over visual objectsShanka Subhra Mondal, Taylor Whittington Webb, Jonathan CohenICLR 2023 · 7 citations
- On The Specialization of Neural ModulesDevon Jarvis, Richard Klein, Benjamin Rosman, Andrew M. SaxeICLR 2023 · 4 citations
- Scaling can lead to compositional generalizationFlorian Redhardt, Yassir Akram, Simon SchugNeurIPS 2025 · 11 citations
- On the generalization capacity of neural networks during generic multimodal reasoningTakuya Ito, Soham Dan, Mattia Rigotti, James R. Kozloski et al.ICLR 2024 · 4 citations
