Deep Learning meets Projective Clustering
Alaa Maalouf, Harry Lang, Daniela Rus, Dan Feldman
Abstract
A common approach for compressing NLP networks is to encode the embedding layer as a matrix , compute its rank- approximation via SVD, and then factor into a pair of matrices that correspond to smaller fully-connected layers to replace the original embedding layer. Geometrically, the rows of represent points in , and the rows of represent their projections onto the -dimensional subspace that minimizes the sum of squared distances ("errors") to the points. In practice, these rows of may be spread around subspaces, so factoring based on a single subspace may lead to large errors that turn into large drops in accuracy. Inspired by projective clustering from computational geometry, we suggest replacing this subspace by a set of subspaces, each of dimension , that minimizes the sum of squared distances over every point (row in ) to its closest subspace. Based on this approach, we provide a novel architecture that replaces the original embedding layer by a set of small layers that operate in parallel and are then recombined with a single fully-connected layer. Extensive experimental results on the GLUE benchmark yield networks that are both more accurate and smaller compared to the standard matrix factorization (SVD). For example, we further compress DistilBERT by reducing the size of the embedding layer by while incurring only a average drop in accuracy over all nine GLUE tasks, compared to a drop using the existing SVD approach. On RoBERTa we achieve compression of the embedding layer with less than a average drop in accuracy as compared to a drop previously. Open code for reproducing and extending our results is provided.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Compressing Neural Networks: Towards Determining the Optimal Layer-wise DecompositionLucas Liebenwein, Alaa Maalouf, Dan Feldman, Daniela RusNeurIPS 2021 · 60 citations
- Pruning Neural Networks via Coresets and Convex Geometry: Towards No AssumptionsMurad Tukan, Loay Mualem, Alaa MaaloufNeurIPS 2022 · 29 citations
- Sparse Flows: Pruning Continuous-depth ModelsLucas Liebenwein, Ramin M. Hasani, Alexander Amini, Daniela RusNeurIPS 2021 · 21 citations
- Model Preserving Compression for Neural NetworksJerry Chee, Megan Flynn, Anil Damle, Christopher De SaNeurIPS 2022 · 19 citations
- Compress to Impress: Efficient LLM Adaptation Using a Single Gradient Step on 100 SamplesShiva Sreeram, Alaa Maalouf, Pratyusha Sharma, Daniela RusNeurIPS 2025 · 2 citations
Builds on4
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Q-BERT: Hessian Based Ultra Low Precision Quantization of BERTSheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma et al.AAAI 2020 · 656 citations
- Structured Pruning of Large Language ModelsZiheng Wang, Jeremy Wohlwend, Tao LeiEMNLP 2020 · 88 citations
- Coresets for Near-Convex FunctionsMurad Tukan, Alaa Maalouf, Dan FeldmanNeurIPS 2020 · 49 citations
Related papers
- DRONE: Data-aware Low-rank Compression for Large NLP ModelsPatrick H. Chen, Hsiang-Fu Yu, Inderjit S. Dhillon, Cho-Jui HsiehNeurIPS 2021 · 109 citations
- ROSITA: Refined BERT cOmpreSsion with InTegrAted techniquesYuanxin Liu, Zheng Lin, Fengcheng YuanAAAI 2021 · 22 citations
- DKM: Differentiable k-Means Clustering Layer for Neural Network CompressionMinsik Cho, Keivan Alizadeh-Vahid, Saurabh Adya, Mohammad RastegariICLR 2022 · 39 citations
- Exploring extreme parameter compression for pre-trained language modelsBenyou Wang, Yuxin Ren, Lifeng Shang, Xin Jiang et al.ICLR 2022 · 23 citations
- BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingCanwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei et al.EMNLP 2020 · 168 citations
