PICACHU: Plug-In CGRA Handling Upcoming Nonlinear Operations in LLMs
Jiajun Qin, Tianhua Xia, Cheng Tan, Jeff Zhang, Sai Qian Zhang
Abstract
Large language models (LLMs) have revolutionized natural language processing (NLP) domain by achieving state-ofthe-art performance across a range of benchmarks. However, nonlinear operations in LLMs significantly contribute to inference latency and present unique challenges that have not been encountered previously. Addressing these challenges requires accelerators that combine efficiency, flexibility, and support for user-defined precision. Our analysis reveals that Coarse-Grained Reconfigurable Arrays (CGRAs) provide an effective solution, offering a balance of performance and flexibility tailored to domain-specific workloads.
This paper introduces PICACHU, a plug-in coarse-grained reconfigurable accelerator tailored to efficiently handle nonlinear operations by using custom algorithms and a dedicated compiler toolchain. PICACHU is the first to target all nonlinear operations within LLMs and to consider CGRA as a plug-in accelerator for LLM inference. Our evaluation shows that PICACHU achieves speedups of 1.86× and 1.55× over prior state-of-the-art accelerators in LLM inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26f4d09b-74db-4f73-83c5-3f3f3236eb2eCited by top-tier papers3
- Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge ComputingTianhua Xia, Sai Qian ZhangMICRO 2025 · 2 citations
- Combating the Memory Walls: Optimization Pathways for Long-Context Agentic Llm InferenceHaoran Wu, Can Xiao, Jiayi Nie, Xuan Guo et al.ISCA 2026
- NEURA: A Unified and Retargetable Compilation Framework for Coarse-Grained Reconfigurable ArchitecturesShangkun Li, Jinming Ge, Diyuan Tao, Zeyu Li et al.PLDI 2026
Builds on51
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- ML-CGRA: An Integrated Compilation Framework to Enable Efficient Machine Learning Acceleration on CGRAsYixuan Luo, Cheng Tan, Nicolas Bohm Agostini, Ang Li et al.DAC 2023 · 39 citations
- Pushing the Limits of BFP on Narrow Precision LLM InferenceHui Wang, Yuan Cheng, Xiaomeng Han, Zhengpeng Zhao et al.AAAI 2025 · 1 citation
- VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible AcceleratorZhican Wang, Hongxiang Fan, Haroon Waris, Gang Wang et al.DAC 2025 · 1 citation
- MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language ModelsElias Frantar, Roberto L. Castro, Jiale Chen, Torsten Hoefler et al.PPoPP 2025 · 24 citations
- NLI : Non-uniform Linear Interpolation Approximation of Nonlinear Operations for Efficient LLMs InferenceJiangyong Yu, Xiaomeng Han, Xing Hu, Chen Xu et al.ICLR 2026 · 2 citations
