Mind the Gap: Removing the Discretization Gap in Differentiable Logic Gate Networks
Shakir Yousefi, Andreas Plesner, Till Aczel, Roger Wattenhofer
摘要
Modern neural networks exhibit state-of-the-art performance on many existing benchmarks, but their high computational requirements and energy usage cause researchers to explore more efficient solutions for real-world deployment. Differentiable logic gate networks (DLGNs) learns a large network of logic gates for efficient image classification. However, learning a network that can solve simple problems like CIFAR-10 or CIFAR-100 can take days to weeks to train. Even then, almost half of the neurons remains unused, causing a discretization gap. This discretization gap hinders real-world deployment of DLGNs, as the performance drop between training and inference negatively impacts accuracy. We inject Gumbel noise with a straight-through estimator during training to significantly speed up training, improve neuron utilization, and decrease the discretization gap. We theoretically show that this results from implicit Hessian regularization, which improves the convergence properties of DLGNs. We train networks 4.5× faster in wall-clock time, reduce the discretization gap by 98%, and reduce the number of unused gates by 100%. * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Light Differentiable Logic Gate NetworksLukas Rüttgers, Till Aczel, Andreas Plesner, Roger WattenhoferICLR 2026 · 被引用 11 次
- Align Forward, Adapt Backward: Closing the Discretization Gap in Logic Gate NetworksYoungsung KimICML 2026 · 被引用 1 次
- Differentiable Weightless Controllers: Learning Logic Circuits for Continuous ControlFabian Kresse, Christoph LampertICML 2026 · 被引用 1 次
- Two-Stage Unit Tying for Simplifying Differentiable Logic Gate NetworksSeungheon Lee, Jeongmin Sun, Jaeyong ChungICML 2026
它引用的顶会 Paper20
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Stabilizing Differentiable Architecture Search via Perturbation-based RegularizationXiangning Chen, Cho-Jui HsiehICML 2020 · 被引用 235 次
相关 Paper
- Deep Differentiable Logic Gate NetworksFelix Petersen, Christian Borgelt, Hilde Kuehne, Oliver DeussenNeurIPS 2022 · 被引用 117 次
- Convolutional Differentiable Logic Gate NetworksFelix Petersen, Hilde Kuehne, Christian Borgelt, Julian Welzel 等NeurIPS 2024 · 被引用 58 次
- Training Discrete Deep Generative Models via Gapped Straight-Through EstimatorTing-Han Fan, Ta-Chung Chi, Alexander I. Rudnicky, Peter J. RamadgeICML 2022 · 被引用 9 次
- Discrete Model Compression With Resource Constraint for Deep Neural NetworksShangqian Gao, Feihu Huang, Jian Pei, Heng HuangCVPR 2020
- Injecting Logical Constraints into Neural Networks via Straight-Through EstimatorsZhun Yang, Joohyung Lee, Chiyoun ParkICML 2022 · 被引用 26 次
