Multiplicative Interactions and Where to Find Them
Siddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz, Jack W. Rae, Simon Osindero, Yee Whye Teh, Tim Harley, Razvan Pascanu
摘要
We explore the role of multiplicative interaction as a unifying framework to describe a range of classical and modern neural network architectural motifs, such as gating, attention layers, hypernetworks, and dynamic convolutions amongst others. Multiplicative interaction layers as primitive operations have a long-established presence in the literature, though this often not emphasized and thus under-appreciated. We begin by showing that such layers strictly enrich the representable function classes of neural networks. We conjecture that multiplicative interactions offer a particularly powerful inductive bias when fusing multiple streams of information or when conditional computation is required. We therefore argue that they should be considered in many situation where multiple compute or information paths need to be combined, in place of the simple and oft-used concatenation operation. Finally, we back up our claims and demonstrate the potential of multiplicative interactions by applying them in large-scale complex RL and sequence modelling tasks, where their use allows us to deliver state-of-the-art results, and thereby provides new evidence in support of multiplicative interactions playing a more prominent role when designing new neural network architectures.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper43
- AutoML-Zero: Evolving Machine Learning Algorithms From ScratchEsteban Real, Chen Liang, David R. So, Quoc V. LeICML 2020 · 被引用 265 次
- Searching for Efficient Transformers for Language ModelingDavid R. So, Wojciech Manke, Hanxiao Liu, Zihang Dai 等NeurIPS 2021 · 被引用 205 次
- Quantifying & Modeling Multimodal Interactions: An Information Decomposition FrameworkPaul Pu Liang, Yun Cheng, Xiang Fan, Chun Kai Ling 等NeurIPS 2023 · 被引用 120 次
- Hit and Lead Discovery with Explorative RL and Fragment-based Molecule GenerationSoojung Yang, Doyeong Hwang, Seul Lee, Seongok Ryu 等NeurIPS 2021 · 被引用 106 次
- Pareto Set Learning for Neural Multi-Objective Combinatorial OptimizationXi Lin, Zhiyuan Yang, Qingfu ZhangICLR 2022 · 被引用 105 次
相关 Paper
- Is Kernel Prediction More Powerful than Gating in Convolutional Neural Networks?Lorenz K. MüllerICML 2024
- HyperMLP: An Integrated Perspective for Sequence ModelingJiecheng Lu, Shihao YangICML 2026 · 被引用 2 次
- A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, GeneralizationMuhammed Ustaomeroglu, Guannan QuICML 2025
- Not All Attention Is Needed: Gated Attention Network for Sequence DataLanqing Xue, Xiaopeng Li, Nevin L. ZhangAAAI 2020 · 被引用 47 次
- Recomposing the Reinforcement Learning Building Blocks with HypernetworksElad Sarafian, Shai Keynan, Sarit KrausICML 2021 · 被引用 42 次
