LoKA: Low-Precision Kernel Applications for Recommendation Models at Scale
Liang Luo, Yinbin Ma, Quanyu Zhu, Vasiliy Kuznetsov, Yuxin Chen, Jian Jiao, Jiecao Yu, Buyun Zhang, Tongyi Tang, Xiaohan Wei, Yanli Zhao, Zeliang Chen
摘要
Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8. While successfully applied to large language models (LLMs), its adoption in large recommendation models (LRMs) has been limited. This is because LRMs are numerically sensitive, dominated by small matrix multiplications (GEMMs) followed by normalization, and trained in communication-intensive environments. Applying FP8 directly to LRMs often degrades model quality and prolongs training time. These challenges are inherent to LRM workloads and cannot be resolved merely by introducing better FP8 kernels. Instead, a system-model co-design approach is needed to successfully integrate FP8. We present LoKA (Low-precision Kernel Applications), a framework that makes FP8 practical for LRMs through three principles: profile under realistic distributions to know where low precision is safe, co-design model components with hardware to expand where it is safe, and orchestrate across kernel libraries to maximize the gains. Concretely, LoKA Probe is a statistically grounded, online benchmarking method that learns activation and weight statistics, and quantifies per-layer errors. This process pinpoints safe and unsafe, fast and slow sites for FP8 adoption. LoKA Mods is a set of reusable model adaptations that improve both numerical stability and execution efficiency with FP8. LoKA Dispatch is a runtime that leverages the statistical insights from LoKA Probe to select the fastest FP8 kernel that satisfies the accuracy requirements. Deployed on production LRMs that serve billions of users with advertising recommendations at a major social media company, LoKA delivers up to 20% higher training throughput and 40% faster inference on heterogeneous GPUs (H100, B200, GB200, MI300X and MI350X) in production environment, with no quality loss, turning FP8 from a modelquality risk into a reliable performance lever for LRMs at scale.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain 等WWW 2021 · 被引用 793 次
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li 等NeurIPS 2024 · 被引用 723 次
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
相关 Paper
- Towards Fully FP8 GEMM LLM Training at ScaleAlejandro Hernández-Cano, Dhia Garbaya, Imanol Schlag, Martin JaggiNeurIPS 2025 · 被引用 13 次
- Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUsRishabh Jain, Vivek M. Bhasi, Adwait Jog, Anand Sivasubramaniam 等MICRO 2024 · 被引用 5 次
- VeriLocc: End-to-End Cross-Architecture Register Allocation via LLMLesheng Jin, Zhenyuan Ruan, Haohui Mai, Jingbo ShangEMNLP 2025
- FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point ArithmeticKanghyun Choi, Hyeyoon Lee, Sunjong Park, Dain Kwon 等NeurIPS 2025 · 被引用 1 次
- Loquetier: A Virtualized Multi-LoRA Framework for Unified LLM Fine-tuning and ServingYuchen Zhang, Hanyue Du, Chun Cao, Jingwei XuNeurIPS 2025 · 被引用 1 次
