oFFN: Outlier and Neuron-aware Structured FFN for Fast yet Accurate LLM Inference
Geunsoo Song, Hoeseok Yang, Youngmin Yi
Abstract
With the advent of large-scale language models (LLMs), various optimization techniques have been proposed to enable efficient inference. Among these, methods that aggressively exploit output activation sparsity have attracted significant attention, which leverage ReLU-fied LLMs and skip the entire memory accesses as well as the computation for the output element if it was predicted as sparse. Achieving fast and accurate prediction of output activation sparsity is crucial to enhancing inference efficiency. However, in practice, phenomena such as activation outliers and hot and cold neurons, which significantly affect the exploitation of sparsity during LLM inference, have either been addressed individually or not structurally integrated in existing work.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b0ee4a2d-dcfe-482d-a4a5-421ff942c98eRelated papers
- Grasp: Group-based Prediction of Activation Sparsity for Fast LLM InferenceJiho Shin, Hoeseok Yang, Youngmin YiDAC 2025
- R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM InferenceZhenyu Zhang, Zechun Liu, Yuandong Tian, Harshit Khaitan et al.ICLR 2025
- ReLU Strikes Back: Exploiting Activation Sparsity in Large Language ModelsIman Mirzadeh, Keivan Alizadeh-Vahid, Sachin Mehta, Carlo C. del Mundo et al.ICLR 2024 · 109 citations
- Universal Properties of Activation Sparsity in Modern Large Language ModelsFilip Szatkowski, Patryk Będkowski, Alessio Devoto, Jan Dubiński et al.ICLR 2026 · 5 citations
- Learn To be Efficient: Build Structured Sparsity in Large Language ModelsHaizhong Zheng, Xiaoyan Bai, Xueshen Liu, Zhuoqing Morley Mao et al.NeurIPS 2024 · 29 citations
