Discrete-Valued Neural Communication
Dianbo Liu, Alex Lamb, Kenji Kawaguchi, Anirudh Goyal, Chen Sun, Michael C. Mozer, Yoshua Bengio
Abstract
Deep learning has advanced from fully connected architectures to structured models organized into components, e.g., the transformer composed of positional elements, modular architectures divided into slots, and graph neural nets made up of nodes. In structured models, an interesting question is how to conduct dynamic and possibly sparse communication among the separate components. Here, we explore the hypothesis that restricting the transmitted information among components to discrete representations is a beneficial bottleneck. The motivating intuition is human language in which communication occurs through discrete symbols. Even though individuals have different understandings of what a"cat"is based on their specific experiences, the shared discrete token makes it possible for communication among individuals to be unimpeded by individual differences in internal representation. To discretize the values of concepts dynamically communicated among specialist components, we extend the quantization mechanism from the Vector-Quantized Variational Autoencoder to multi-headed discretization with shared codebooks and use it for discrete-valued neural communication (DVNC). Our experiments show that DVNC substantially improves systematic generalization in a variety of architectures -- transformers, modular architectures, and graph neural networks. We also show that the DVNC is robust to the choice of hyperparameters, making the method very useful in practice. Moreover, we establish a theoretical justification of our discretization process, proving that it has the ability to increase noise robustness and reduce the underlying dimensionality of the model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Universal Humanoid Motion Representations for Physics-Based ControlZhengyi Luo, Jinkun Cao, Josh Merel, Alexander Winkler et al.ICLR 2024 · 125 citations
- How Does Information Bottleneck Help Deep Learning?Kenji Kawaguchi, Zhun Deng, Xu Ji, Jiaoyang HuangICML 2023 · 117 citations
- Neural Systematic BinderGautam Singh, Yeongbin Kim, Sungjin AhnICLR 2023 · 105 citations
- Disentanglement via Latent QuantizationKyle Hsu, William Dorrell, James C. R. Whittington, Jiajun Wu et al.NeurIPS 2023 · 54 citations
- Learning Invariant Molecular Representation in Latent Discrete SpaceXiang Zhuang, Qiang Zhang, Keyan Ding, Yatao Bian et al.NeurIPS 2023 · 41 citations
Builds on4
- Recurrent Independent MechanismsAnirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani et al.ICLR 2021 · 357 citations
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 322 citations
- Coordination Among Neural Modules Through a Shared Global WorkspaceAnirudh Goyal, Aniket Rajiv Didolkar, Alex Lamb, Kartikeya Badola et al.ICLR 2022 · 114 citations
- Neural Production SystemsAniket Didolkar, Anirudh Goyal, Nan Rosemary Ke, Charles Blundell et al.NeurIPS 2021 · 94 citations
Related papers
- Adaptive Discrete Communication Bottlenecks with Dynamic Vector Quantization for Heterogeneous Representational CoarsenessDianbo Liu, Alex Lamb, Xu Ji, Pascal Tikeng Notsawo Jr. et al.AAAI 2023 · 19 citations
- MoToRec: Sparse-Regularized Multimodal Tokenization for Cold-Start RecommenderJialin Liu, Zhaorui Zhang, Ray C. C. CheungAAAI 2026
- Learning Graph Quantized TokenizersLimei Wang, Kaveh Hassani, Si Zhang, Dongqi Fu et al.ICLR 2025
- Neural Language of Thought ModelsYi-Fu Wu, Minseung Lee, Sungjin AhnICLR 2024 · 11 citations
- VQGraph: Rethinking Graph Representation Space for Bridging GNNs and MLPsLing Yang, Ye Tian, Minkai Xu, Zhongyi Liu et al.ICLR 2024 · 48 citations
