Efficient Learning of Discrete-Continuous Computation Graphs
David Friede, Mathias Niepert
Abstract
Numerous models for supervised and reinforcement learning benefit from combinations of discrete and continuous model components. End-to-end learnable discrete-continuous models are compositional, tend to generalize better, and are more interpretable. A popular approach to building discrete-continuous computation graphs is that of integrating discrete probability distributions into neural networks using stochastic softmax tricks. Prior work has mainly focused on computation graphs with a single discrete component on each of the graph's execution paths. We analyze the behavior of more complex stochastic computations graphs with multiple sequential discrete components. We show that it is challenging to optimize the parameters of these models, mainly due to small gradients and local minima. We then propose two new strategies to overcome these challenges. First, we show that increasing the scale parameter of the Gumbel noise perturbations during training improves the learning behavior. Second, we propose dropout residual connections specifically tailored to stochastic, discrete-continuous computation graphs. With an extensive set of experiments, we show that we can train complex discrete-continuous models which one cannot train with standard stochastic softmax tricks. We also show that complex discrete-stochastic models generalize better than their continuous counterparts on several benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d91e9cb-bb7b-43a9-8f3f-a4732e482fe9Cited by top-tier papers1
Ask how each one uses itBuilds on6
- You CAN Teach an Old Dog New Tricks! On Training Knowledge Graph EmbeddingsDaniel Ruffinelli, Samuel Broscheit, Rainer GemullaICLR 2020 · 238 citations
- Gradient Estimation with Stochastic Softmax TricksMax B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause et al.NeurIPS 2020 · 104 citations
- Neural-Symbolic Integration: A Compositional PerspectiveEfthymia Tsamoura, Timothy M. Hospedales, Loizos MichaelAAAI 2021 · 85 citations
- Efficient Marginalization of Discrete and Structured Latent Variables via SparsityGonçalo M. Correia, Vlad Niculae, Wilker Aziz, André F. T. MartinsNeurIPS 2020 · 25 citations
- LP-SparseMAP: Differentiable Relaxed Optimization for Sparse Structured PredictionVlad Niculae, André F. T. MartinsICML 2020 · 22 citations
Related papers
- Invertible Gaussian Reparameterization: Revisiting the Gumbel-SoftmaxAndres Potapczynski, Gabriel Loaiza-Ganem, John P. CunninghamNeurIPS 2020 · 44 citations
- Training Discrete Deep Generative Models via Gapped Straight-Through EstimatorTing-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. RamadgeICML 2022 · 9 citations
- Gradient-Based Program Synthesis with Neurally Interpreted LanguagesMatthew Macfarlane, Clément Bonnet, Herke van Hoof, Levi LelisICLR 2026 · 3 citations
- LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft ThinkingJunhong Wu, Jinliang Lu, Zixuan Ren, Gangqiang Hu et al.ICLR 2026 · 35 citations
- Dynamic Convolutions: Exploiting Spatial Sparsity for Faster InferenceThomas Verelst, Tinne TuytelaarsCVPR 2020
