HASTE: Hardware-Aware Dynamic Sparse Training for Large Output Spaces
Nasib Ullah, Jinbin Zhang, Jean Lucien Randrianantenaina, Erik Schultheis, Rohit Babbar
摘要
Extreme multi-label classification (XMC) involves learning models over large output spaces with millions of labels, making the output layer a memory-compute bottleneck. While sparsity-based methods reduce arithmetic complexity, they often fail to yield proportional speedups due to irregular memory access, poor hardware utilization, or reliance on auxiliary architectural components in long-tailed regimes. We introduce group-shared fixed fan-in sparsity, a semi-structured output-layer design in which semantically related labels share a sparse input pattern while retaining independent weights. This grouping introduces a task-aligned inductive bias---encouraging related labels to share feature subsets---while reducing index memory overhead, increasing feature reuse across labels, and enabling efficient GPU execution via custom CUDA kernels that leverage modern accelerator primitives. As an alternative to auxiliary objectives, we exploit the long-tailed structure of XMC by decomposing the output layer into a small dense head over frequent labels and a group-shared sparse tail over the remainder, providing an informative gradient pathway while preserving the memory benefits of sparsity. Through kernel-level microbenchmarking, we show that group-shared fixed fan-in translates arithmetic reductions into practical wall-clock gains, achieving up to speedup in the forward pass and up to speedup in backward passes over standard fixed fan-in sparsity, while operating within a few percent of a FLOPs-matched dense bottleneck. Across large-scale XMC benchmarks, our approach matches or improves precision@k over prior sparse baselines, while narrowing the performance gap to dense.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu 等ICLR 2021 · 被引用 301 次
- Fast Multi-Resolution Transformer Fine-tuning for Extreme Multi-label Text ClassificationJiong Zhang, Wei-Cheng Chang, Hsiang-Fu Yu, Inderjit S. DhillonNeurIPS 2021 · 被引用 147 次
- SiameseXML: Siamese Networks meet Extreme Classifiers with 100M LabelsKunal Dahiya, Ananye Agarwal, Deepak Saini, Gururaj K 等ICML 2021 · 被引用 61 次
- Dynamic Sparse Training with Structured SparsityMike Lasby, Anna Golubeva, Utku Evci, Mihai Nica 等ICLR 2024 · 被引用 37 次
相关 Paper
- Towards Robust Prediction on Tail LabelsTong Wei, Wei-Wei Tu, Yufeng Li, Guo-Ping YangKDD 2021 · 被引用 12 次
- Navigating Extremes: Dynamic Sparsity in Large Output SpacesNasibullah Nasibullah, Erik Schultheis, Mike Lasby, Yani Ioannou 等NeurIPS 2024
- Generalized test utilities for long-tail performance in extreme multi-label classificationErik Schultheis, Marek Wydmuch, Wojciech Kotlowski, Rohit Babbar 等NeurIPS 2023 · 被引用 7 次
- ELIAS: End-to-End Learning to Index and Search in Large Output SpacesNilesh Gupta, Patrick H. Chen, Hsiang-Fu Yu, Cho-Jui Hsieh 等NeurIPS 2022 · 被引用 19 次
- Extreme Multi-label Classification from Aggregated LabelsYanyao Shen, Hsiang-Fu Yu, Sujay Sanghavi, Inderjit S. DhillonICML 2020 · 被引用 10 次
