ANN Softmax: Acceleration of Extreme Classification Training
Kang Zhao, Liuyihan Song, Yingya Zhang, Pan Pan, Yinghui Xu, Rong Jin
摘要
Thanks to the popularity of GPU and the growth of its computational power, more and more deep learning tasks, such as face recognition, image retrieval and word embedding, can take advantage of extreme classification to improve accuracy. However, it remains a big challenge to train a deep model with millions of classes efficiently due to the huge memory and computation consumption in the last layer. By sampling a small set of classes to avoid the total classes calculation, sampling-based approaches have been proved to be an effective solution. But most of them suffer from the following two issues: i) the important classes are ignored or only partly sampled, such as the methods using random sampling scheme or retrieval techniques of low recall (e.g., locality-sensitive hashing), resulting in the degradation of accuracy; ii) inefficient implementation owing to incompatibility with GPU, like selective softmax. It uses hashing forest to help select classes, but the search process is implemented in CPU. To address the above issues, we propose a new sampling-based softmax called ANN Softmax in this paper. Specifically, we employ binary quantization with inverted file system to improve the recall of important classes. With the help of dedicated kernel design, it can be totally parallelized in mainstream training framework. Then, we find the size of important classes that are recalled by each training sample has a great impact on the final accuracy, so we introduce sample grouping optimization to well approximate the full classes training. Experimental evaluations on two tasks (Embedding Learning and Classification) and ten datasets (e.g., MegaFace, ImageNet, SKU datasets) demonstrate our proposed method maintains the same precision as Full Softmax for different loss objectives, including cross entropy loss, ArcFace, CosFace and D-Softmax loss, with only 1/10 sampled classes, which outperforms the state-of-the-art techniques. Moreover, we implement ANN Soft-max in a complete GPU pipeline that can accelerate the training more than 4.3X. Equipped our method with a 256 GPUs cluster, the time of training a classifier of 300 million classes on our SKU-300M dataset can be reduced to ten days.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- Softmax Dissection: Towards Understanding Intra- and Inter-Class Objective for Embedding LearningLanqing He, Zhongdao Wang, Yali Li, Shengjin WangAAAI 2020 · 被引用 34 次
- Panorama: A Data System for Unbounded Vocabulary Querying over VideoYuhao Zhang, Arun KumarVLDB 2020 · 被引用 28 次
- Deep or Simple Models for Semantic Tagging? It Depends on your DataJinfeng Li, Yuliang Li, Xiaolan Wang, Wang-Chiew TanVLDB 2020
- ODIN: Automated Drift Detection and Recovery in Video AnalyticsAbhijit Suprem, Joy Arulraj, Calton Pu, João Eduardo FerreiraVLDB 2020
相关 Paper
- Extreme Classification via Adversarial Softmax ApproximationRobert Bamler, Stephan MandtICLR 2020 · 被引用 25 次
- A Tale of Two Efficient and Informative Negative Sampling DistributionsShabnam Daghaghi, Tharun Medini, Nicholas Meisburger, Beidi Chen 等ICML 2021 · 被引用 11 次
- Unleashing the Full Potential of Product Quantization for Large-Scale Image RetrievalYu Liang, Shiliang Zhang, Li Ken Li, Xiaoyu WangNeurIPS 2023 · 被引用 5 次
- ECSSD: Hardware/Data Layout Co-Designed In-Storage-Computing Architecture for Extreme ClassificationSiqi Li, Fengbin Tu, Liu Liu, Jilan Lin 等ISCA 2023 · 被引用 17 次
- Softmax Tree: An Accurate, Fast Classifier When the Number of Classes Is LargeArman Zharmagambetov, Magzhan Gabidolla, Miguel Á. Carreira-PerpiñánEMNLP 2021 · 被引用 4 次
