Lune

ICML2026Top-tier venue

Improving the Robustness-Utility Trade-off in Decentralized Learning over Sparse Networks

Yangnan Li, Xuanyu Cao, Shenghui Song

2026Year

Abstract

Resilience against Byzantine attackers and faster convergence on sparse networks are critical for decentralized optimization, yet existing methods fail to achieve both simultaneously. Existing DSGD-based Byzantine-resilient methods suffer from high transient complexity of O((1−λ)−6)\mathcal{O}\left((1-\lambda)^{-6}\right), where 1−λ1-\lambda denotes the spectral gap of the network. While bias-correction methods such as Exact Diffusion can improve topology dependence, directly combining them with robust aggregators can lead to error accumulation. To address this issue, we introduce the scaled dual ascent (SDA) within the augmented Lagrangian framework for decentralized optimization, which mitigates error accumulation by scaling the dual update steps. Based on this, we propose BRED, which integrates Byzantine-robust Exact Diffusion with the SDA framework. We prove that BRED attains linear speedup, and achieves transient complexity of O((1−λ)−2)\mathcal{O}\left((1-\lambda)^{-2}\right) when the Byzantine fraction δ\delta is small. We further propose the momentum variant BRED-M, which reduces the Byzantine-affected transient complexity from O(δ2(1−λ)−6)\mathcal{O}\left(\delta^2(1-\lambda)^{-6}\right) to O(δ2(1−λ)−4)\mathcal{O}\left(\delta^2(1-\lambda)^{-4}\right). Empirical results on benchmark datasets demonstrate the efficacy of the proposed methods across diverse network topologies.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 9c19b591-b895-4b49-b416-371934f47f78

Builds on4

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines