Bringing Near Data Processing Into the Low-Bit Floating-Point Era
Tongxin Xie, Mingyu Gao, Zehao Wang, Zhihao Jia, Yuechen Xi, Bing Li, Mo Guang, Jiale Yan, Kaiwen Long, Xingcheng Zhang, Huazhong Yang, Yuan Xie
摘要
Near data processing (NDP) based on DRAM has emerged to be a promising solution to the “memory wall” problem of machine learning models. From the algorithmic perspective, group-wise low-bit floating-point (FP) quantization has become an important trend for both efficient training and inference. Integrating low-bit FP quantization into NDP also notably shrinks the memory footprint of large models, alleviating the memory-capacity constraints of NDP architectures. However, existing NDP compilers struggle to support efficient low-bit FP computation on NDP. First, different quantization configurations exhibit different preferences for NDP compilation strategies. Second, the fine-grained grouping leads to frequent switching between quantized value access and group scale access during computation, increasing the DRAM row-buffer miss rate. Third, fine-grained grouping triggers frequent high-precision dequantization operations, causing significant latency overhead. To address these challenges, this paper proposes FlexQ-NDP, an NDP compiler tailored for general low-bit FP computation. Firstly, we develop an open-source simulation framework 11 Available at https://github.com/ISCA26-FlexQ-NDP-ae/flexqndp to model the low-bit FP computation overhead on NDP. Secondly, we design a scale-value interleaved FP layout, effectively reducing DRAM row-changing overhead. Thirdly, we propose a dequantization-hiding technique based on instruction reordering to reduce the DRAM idle time induced by frequent dequantization operations. Finally, we develop a lightweight compilationspace pruning and search strategy to enable efficient low-bit FP computation on NDP. Extensive experiments show that FlexQ-NDP achieves up to speedup over existing compilation strategies on various low-bit FP quantization configurations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li 等NeurIPS 2024 · 被引用 723 次
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks 等ISCA 2020 · 被引用 235 次
- Newton: A DRAM-maker's Accelerator-in-Memory (AiM) Architecture for Machine LearningMingxuan He, Choungki Song, Ilkon Kim, Chunseok Jeong 等MICRO 2020 · 被引用 208 次
- Pushing the Limits of Narrow Precision Inferencing at Cloud Scale with Microsoft Floating PointBita Darvish Rouhani, Daniel Lo, Ritchie Zhao, Ming Liu 等NeurIPS 2020 · 被引用 153 次
相关 Paper
- UniNDP: A Unified Compilation and Simulation Tool for Near DRAM Processing ArchitecturesTongxin Xie, Zhenhua Zhu, Bing Li, Yukai He 等HPCA 2025 · 被引用 9 次
- LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language ModelsGunho Park, Baeseong Park, Minsub Kim, Sungjae Lee 等ICLR 2024 · 被引用 134 次
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li 等ICCV 2019 · 被引用 540 次
- TRiM: Enhancing Processor-Memory Interfaces with Scalable Tensor Reduction in MemoryJaehyun Park, Byeongho Kim, Sungmin Yun, Eojin Lee 等MICRO 2021 · 被引用 70 次
- FlexiQ: Adaptive Mixed-Precision Quantization for Latency/Accuracy Trade-Offs in Deep Neural NetworksJaemin Kim, Hongjun Um, Sungkyun Kim, Yongjun Park 等EuroSys 2026
