Lune

ICML2026Top-tier venue

SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs

Mikołaj Zasada, Łukasz Struski, Jacek Tabor, Marcin Kurdziel

2026Year
2Citations

Abstract

Sparse Mixture-of-Experts (MoE) architectures enable scaling LLM parameters under a fixed inference budget by activating only a small subset of experts via top-kk routing. While this preserves causality and suits autoregressive language models, the discrete top-kk operator is not differentiable, forcing a fixed number of active experts per input and resulting in inefficient use of computation. We propose SoftMoE, which replaces discrete routing with a truncated soft top-kk LapSum relaxation, allowing gradient-based optimization of expert routing. We further parameterize the mean number of active experts per layer and impose a global budget constraint, enabling the model to learn how to allocate expert capacity across layers. SoftMoE remains fully compatible with autoregressive modeling and achieves performance comparable to or better than sparse MoE on language modeling and downstream tasks, while activating significantly fewer experts. Notably, the learned allocation is highly non-uniform, with later layers activating more experts. The source code is publicly available†^\dagger.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 5fd96c29-3d2b-46a1-b490-bdcdbcfc3b14

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines