Lune

ASPLOS2026顶会

oFFN: Outlier and Neuron-aware Structured FFN for Fast yet Accurate LLM Inference

Geunsoo Song, Hoeseok Yang, Youngmin Yi

2026年份

摘要

With the advent of large-scale language models (LLMs), various optimization techniques have been proposed to enable efficient inference. Among these, methods that aggressively exploit output activation sparsity have attracted significant attention, which leverage ReLU-fied LLMs and skip the entire memory accesses as well as the computation for the output element if it was predicted as sparse. Achieving fast and accurate prediction of output activation sparsity is crucial to enhancing inference efficiency. However, in practice, phenomena such as activation outliers and hot and cold neurons, which significantly affect the exploitation of sparsity during LLM inference, have either been addressed individually or not structurally integrated in existing work.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

lune papers get b0ee4a2d-dcfe-482d-a4a5-421ff942c98e

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖