Lune

NeurIPS2025顶会

Differentiable Sparsity via DD-Gating: Simple and Versatile Structured Penalization

Chris Kolb, Laetitia Frost, Bernd Bischl, David Rügamer

2025年份
4被引次数

摘要

Structured sparsity regularization offers a principled way to compact neural networks, but its non-differentiability breaks compatibility with conventional stochastic gradient descent and requires either specialized optimizers or additional post-hoc pruning without formal guarantees. In this work, we propose DD-Gating, a fully differentiable structured overparameterization that splits each group of weights into a primary weight vector and multiple scalar gating factors. We prove that any local minimum under DD-Gating is also a local minimum using non-smooth structured L2,2/DL_{2,2/D} penalization, and further show that the DD-Gating objective converges at least exponentially fast to the L2,2/DL_{2,2/D}-regularized loss in the gradient flow limit. Together, our results show that DD-Gating is theoretically equivalent to solving the original group sparsity problem, yet induces distinct learning dynamics that evolve from a non-sparse regime into sparse optimization. We validate our theory across vision, language, and tabular tasks, where DD-Gating consistently delivers strong performance-sparsity tradeoffs and outperforms both direct optimization of structured penalties and conventional pruning baselines.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper21

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖