Neural Attentive Circuits
Martin Weiss, Nasim Rahaman, Francesco Locatello, Chris Pal, Yoshua Bengio, Bernhard Schölkopf, Li Erran Li, Nicolas Ballas
摘要
Recent work has seen the development of general purpose neural architectures that can be trained to perform tasks across diverse data modalities. General purpose models typically make few assumptions about the underlying data-structure and are known to perform well in the large-data regime. At the same time, there has been growing interest in modular neural architectures that represent the data using sparsely interacting modules. These models can be more robust out-ofdistribution, computationally efficient, and capable of sample-efficient adaptation to new data. However, they tend to make domain-specific assumptions about the data, and present challenges in how module behavior (i.e., parameterization) and connectivity (i.e., their layout) can be jointly learned. In this work, we introduce a general purpose, yet modular neural architecture called Neural Attentive Circuits (NACs) that jointly learns the parameterization and a sparse connectivity of neural modules without using domain knowledge. NACs are best understood as the combination of two systems that are jointly trained end-to-end: one that determines the module configuration and the other that executes it on an input. We demonstrate qualitatively that NACs learn diverse and meaningful module configurations on the Natural Language and Visual Reasoning for Real (NLVR2) dataset without additional supervision. Quantitatively, we show that by incorporating modularity in this way, NACs improve upon a strong non-modular baseline in terms of low-shot adaptation on CIFAR and Caltech-UCSD Birds dataset (CUB) by about 10 percent, and OOD robustness on Tiny ImageNet-R by about 2.5 percent. Further, we find that NACs can achieve an 8x speedup at inference time while losing less than 3 percent performance. Finally, we find NACs to yield competitive results on diverse data modalities spanning point-cloud classification, symbolic processing and textclassification from ASCII bytes, thereby confirming its general purpose nature. Introduction General purpose neural models like Perceivers [29] do not make significant assumptions about the underlying data-structure of the input and tend to perform well in the large-data regime. This enables the application of the same model on a variety of data modalities, including images, text, audio, point-clouds, and arbitrary combinations thereof [29, 28] . This is appealing from an ease-of-use perspective, since the amount of domain-specific components is minimized, and the resulting models can function well out-of-the-box in larger machine learning pipelines, e.g., AlphaStar [28] . At the same time, natural data generating processes can often be well-represented by a system of sparsely interacting independent mechanisms [41, 46] , and the Sparse Mechanism Shift hypothesis 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ELDEN: Exploration via Local DependenciesZizhao Wang, Jiaheng Hu, Peter Stone, Roberto Martín-MartínNeurIPS 2023 · 被引用 15 次
- Scalable Modular Network: A Framework for Adaptive Learning via Agreement RoutingMinyang Hu, Hong Chang, Bingpeng Ma, Shiguang Shan 等ICLR 2024 · 被引用 2 次
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals 等ICML 2021 · 被引用 1,399 次
相关 Paper
- Is a Modular Architecture Enough?Sarthak Mittal, Yoshua Bengio, Guillaume LajoieNeurIPS 2022 · 被引用 62 次
- Dynamic Inference with Neural InterpretersNasim Rahaman, Muhammad Waleed Gondal, Shruti Joshi, Peter V. Gehler 等NeurIPS 2021 · 被引用 35 次
- Composable Sparse Subnetworks via Maximum-Entropy PrincipleFrancesco Caso, Samuele Fonio, Simone Monaco, Nicola Saccomanno 等ICLR 2026
- On The Specialization of Neural ModulesDevon Jarvis, Richard Klein, Benjamin Rosman, Andrew M. SaxeICLR 2023 · 被引用 4 次
- Uni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEsJinguo Zhu, Xizhou Zhu, Wenhai Wang, Xiaohua Wang 等NeurIPS 2022 · 被引用 95 次
