Principled Weight Initialisation for Input-Convex Neural Networks
Pieter-Jan Hoedt, Günter Klambauer
Abstract
Input-Convex Neural Networks (ICNNs) are networks that guarantee convexity in their input-output mapping. These networks have been successfully applied for energy-based modelling, optimal transport problems and learning invariances. The convexity of ICNNs is achieved by using non-decreasing convex activation functions and non-negative weights. Because of these peculiarities, previous initialisation strategies, which implicitly assume centred weights, are not effective for ICNNs. By studying signal propagation through layers with non-negative weights, we are able to derive a principled weight initialisation for ICNNs. Concretely, we generalise signal propagation theory by removing the assumption that weights are sampled from a centred distribution. In a set of experiments, we demonstrate that our principled initialisation effectively accelerates learning in ICNNs and leads to better generalisation. Moreover, we find that, in contrast to common belief, ICNNs can be trained without skip-connections when initialised correctly. Finally, we apply ICNNs to a real-world drug discovery task and show that they allow for more effective molecular latent space exploration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Optimal Flow Matching: Learning Straight Trajectories in Just One StepNikita Kornilov, Petr Mokrov, Alexander V. Gasnikov, Alexander KorotinNeurIPS 2024 · 93 citations
- Revisiting Residual Connections: Orthogonal Updates for Stable and Efficient Deep NetworksGiyeong Oh, Woohyun Cho, Siyeol Kim, Suhwan Choi et al.NeurIPS 2025 · 3 citations
- Amortized Maximum Inner Product Search with Learned Support FunctionsTheo X. Olausson, Joao Monteiro, Michal Klein, Marco CuturiICML 2026 · 1 citation
- DilateQuant: Accurate and Efficient Quantization-Aware Training for Diffusion Models via Weight DilationXuewen Liu, Zhikai Li, Minghao Jiang, Mengjuan Chen et al.ACM MM 2025
Builds on5
- Optimal transport mapping via input convex neural networksAshok Vardhan Makkuva, Amirhossein Taghvaei, Sewoong Oh, Jason D. LeeICML 2020 · 254 citations
- Principled Weight Initialization for HypernetworksOscar Chang, Lampros Flokas, Hod LipsonICLR 2020 · 87 citations
- Deep Learning without Shortcuts: Shaping the Kernel with Tailored RectifiersGuodong Zhang, Aleksandar Botev, James MartensICLR 2022 · 30 citations
- Characterizing signal propagation to close the performance gap in unnormalized ResNetsAndrew Brock, Soham De, Samuel L. SmithICLR 2021 · 21 citations
- Non-Gaussian Tensor ProgramsEugene A. Golikov, Greg YangNeurIPS 2022 · 11 citations
Related papers
- Large-Scale Wasserstein Gradient FlowsPetr Mokrov, Alexander Korotin, Lingxiao Li, Aude Genevay et al.NeurIPS 2021 · 112 citations
- AutoInit: Analytic Signal-Preserving Weight Initialization for Neural NetworksGarrett Bingham, Risto MiikkulainenAAAI 2023 · 6 citations
- Canonical Tree Cover Neural Networks for Expressive and Invariant Graph LearningMichael Ito, Danai Koutra, Jenna WiensICLR 2026
- Optimization-Induced Graph Implicit Nonlinear DiffusionQi Chen, Yifei Wang, Yisen Wang, Jiansheng Yang et al.ICML 2022 · 44 citations
- Geometric and Physical Quantities improve E(3) Equivariant Message PassingJohannes Brandstetter, Rob Hesselink, Elise van der Pol, Erik J. Bekkers et al.ICLR 2022 · 307 citations
