Lune

ICLR2024Top-tier venue

SAS: Structured Activation Sparsification

Yusuke Sekikawa, Shingo Yashima

2024Year
1Citations

Abstract

Wide networks usually yield better accuracy than their narrower counterpart at the expense of the massive mult cost. To break this tradeoff, we advocate a novel concept of Structured Activation Sparsification, dubbed SAS, which boosts accuracy without increasing computation by utilizing the projected sparsity in activation maps with a specific structure. Concretely, the projected sparse activation is allowed to have N nonzero value among M consecutive activations. Owing to the local structure in sparsity, the wide matmul between a dense weight and the sparse activation is executed as an equivalent narrow matmul between a dense weight and dense activation, which is compatible with NVIDIA's Sparse Tensor Core developed for the N : M structured sparse weight. In extensive experiments, we demonstrate that increasing sparsity monotonically improves accuracy (up to 7% on CIFAR10) without increasing the mult count. Furthermore, we show that structured sparsification of activation scales better than that of weight given the same computational budget.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 921418e8-c00f-42cc-9c9b-3792f0d70d9f

Builds on9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines