Lune

ASPLOS2026Top-tier venue

oFFN: Outlier and Neuron-aware Structured FFN for Fast yet Accurate LLM Inference

Geunsoo Song, Hoeseok Yang, Youngmin Yi

2026Year

Abstract

With the advent of large-scale language models (LLMs), various optimization techniques have been proposed to enable efficient inference. Among these, methods that aggressively exploit output activation sparsity have attracted significant attention, which leverage ReLU-fied LLMs and skip the entire memory accesses as well as the computation for the output element if it was predicted as sparse. Achieving fast and accurate prediction of output activation sparsity is crucial to enhancing inference efficiency. However, in practice, phenomena such as activation outliers and hot and cold neurons, which significantly affect the exploitation of sparsity during LLM inference, have either been addressed individually or not structurally integrated in existing work.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get b0ee4a2d-dcfe-482d-a4a5-421ff942c98e

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines