Lune

ICLR2026Top-tier venue

FedMuon: Federated Learning with Bias-corrected LMO-based Optimization

Yuki Takezawa, Anastasia Koloskova, Xiaowen Jiang, Sebastian U. Stich

2026Year
9Citations
1Top-tier citations

Abstract

Recently, a new optimization method based on the linear minimization oracle (LMO), called Muon, has been attracting increasing attention since it can train neural networks faster than the existing adaptive optimization methods, such as Adam. In this paper, we study how Muon can be utilized in federated learning. We first show that straightforwardly using Muon as the local optimizer of FedAvg does not work since the LMO is a biased operator. We then propose FedMuon, which can mitigate this issue and can converge to the stationary point. We also analyze how solving the LMO approximately affects the convergence rate and find that, surprisingly, FedMuon can converge for any number of Newton-Schulz iterations, while it can converge faster as we solve the LMO more accurately. Through experiments, we demonstrated that FedMuon can outperform the state-of-the-art federated learning methods.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 92f85e53-56a6-4a1c-8b1e-17af1ee2a81f

Cited by top-tier papers1

Ask how each one uses it

Builds on29

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines