Discrete-Valued Neural Communication
Dianbo Liu, Alex Lamb, Kenji Kawaguchi, Anirudh Goyal, Chen Sun, Michael C. Mozer, Yoshua Bengio
摘要
Deep learning has advanced from fully connected architectures to structured models organized into components, e.g., the transformer composed of positional elements, modular architectures divided into slots, and graph neural nets made up of nodes. In structured models, an interesting question is how to conduct dynamic and possibly sparse communication among the separate components. Here, we explore the hypothesis that restricting the transmitted information among components to discrete representations is a beneficial bottleneck. The motivating intuition is human language in which communication occurs through discrete symbols. Even though individuals have different understandings of what a"cat"is based on their specific experiences, the shared discrete token makes it possible for communication among individuals to be unimpeded by individual differences in internal representation. To discretize the values of concepts dynamically communicated among specialist components, we extend the quantization mechanism from the Vector-Quantized Variational Autoencoder to multi-headed discretization with shared codebooks and use it for discrete-valued neural communication (DVNC). Our experiments show that DVNC substantially improves systematic generalization in a variety of architectures -- transformers, modular architectures, and graph neural networks. We also show that the DVNC is robust to the choice of hyperparameters, making the method very useful in practice. Moreover, we establish a theoretical justification of our discretization process, proving that it has the ability to increase noise robustness and reduce the underlying dimensionality of the model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Universal Humanoid Motion Representations for Physics-Based ControlZhengyi Luo, Jinkun Cao, Josh Merel, Alexander Winkler 等ICLR 2024 · 被引用 125 次
- How Does Information Bottleneck Help Deep Learning?Kenji Kawaguchi, Zhun Deng, Xu Ji, Jiaoyang HuangICML 2023 · 被引用 117 次
- Neural Systematic BinderGautam Singh, Yeongbin Kim, Sungjin AhnICLR 2023 · 被引用 105 次
- Disentanglement via Latent QuantizationKyle Hsu, William Dorrell, James C. R. Whittington, Jiajun Wu 等NeurIPS 2023 · 被引用 54 次
- Learning Invariant Molecular Representation in Latent Discrete SpaceXiang Zhuang, Qiang Zhang, Keyan Ding, Yatao Bian 等NeurIPS 2023 · 被引用 41 次
它引用的顶会 Paper4
- Recurrent Independent MechanismsAnirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani 等ICLR 2021 · 被引用 357 次
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 被引用 322 次
- Coordination Among Neural Modules Through a Shared Global WorkspaceAnirudh Goyal, Aniket Rajiv Didolkar, Alex Lamb, Kartikeya Badola 等ICLR 2022 · 被引用 114 次
- Neural Production SystemsAniket Didolkar, Anirudh Goyal, Nan Rosemary Ke, Charles Blundell 等NeurIPS 2021 · 被引用 94 次
相关 Paper
- Adaptive Discrete Communication Bottlenecks with Dynamic Vector Quantization for Heterogeneous Representational CoarsenessDianbo Liu, Alex Lamb, Xu Ji, Pascal Tikeng Notsawo Jr. 等AAAI 2023 · 被引用 19 次
- MoToRec: Sparse-Regularized Multimodal Tokenization for Cold-Start RecommenderJialin Liu, Zhaorui Zhang, Ray C. C. CheungAAAI 2026
- Learning Graph Quantized TokenizersLimei Wang, Kaveh Hassani, Si Zhang, Dongqi Fu 等ICLR 2025
- Neural Language of Thought ModelsYi-Fu Wu, Minseung Lee, Sungjin AhnICLR 2024 · 被引用 11 次
- VQGraph: Rethinking Graph Representation Space for Bridging GNNs and MLPsLing Yang, Ye Tian, Minkai Xu, Zhongyi Liu 等ICLR 2024 · 被引用 48 次
