Sign-SGD via Parameter-Free Optimization
Daniil Medyakov, Sergey Stanko, Gleb Molodtsov, Philip Zmushko, Grigoriy Evseev, Egor Petrov, Aleksandr Beznosikov
Abstract
Large language models have achieved major advances across domains, yet training them remains extremely resource-intensive. We revisit Sign-SGD, which serves both as a memory-efficient optimizer for single-node training and as a gradient compression mechanism for distributed learning. This paper addresses a central limitation: the effective stepsize cannot be determined a priori because it relies on unknown, problem-specific quantities. We present a parameter-free Sign-SGD that removes manual stepsize selection. We analyze the deterministic single-node case, and extend the method to stochastic single-node training and multi-node settings. We also incorporate the momentum technique into our algorithms and propose a memory-efficient variant that stores only gradient signs instead of full gradients. We evaluate our methods on pre-training LLaMA models (130M and 350M) and fine-tuning a Swin Transformer (28M). Across considered tasks, the proposed methods match the performance of tuned Sign-SGD and AdamW (grid-searched stepsizes with a cosine schedule), while avoiding tuning overhead. Employing parameter-free training yields approximately end-to-end speedup compared to runs with grid-searched stepsizes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a0067df5-0a71-4b1c-add6-cdb4adf82287Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- ReLoRA: High-Rank Training Through Low-Rank UpdatesVladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna RumshiskyICLR 2024 · 214 citations
- Prodigy: An Expeditiously Adaptive Parameter-Free LearnerKonstantin Mishchenko, Aaron DefazioICML 2024 · 131 citations
- Learning-Rate-Free Learning by D-AdaptationAaron Defazio, Konstantin MishchenkoICML 2023 · 117 citations
- DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size ScheduleMaor Ivgi, Oliver Hinder, Yair CarmonICML 2023 · 98 citations
Related papers
- Arbitrary-Order Block SignSGD for Memory-Efficient LLM Fine-TuningYijie Zhou, Shi PuICLR 2026
- SLAMB: Accelerated Large Batch Training with Sparse CommunicationHang Xu, Wenxuan Zhang, Jiawei Fei, Yuzhe Wu et al.ICML 2023 · 7 citations
- SWAN: SGD with Normalization and Whitening Enables Stateless LLM TrainingChao Ma, Wenbo Gong, Meyer Scetbon, Edward MeedsICML 2025
- Stochastic Sign Descent Methods: New Algorithms and Better TheoryMher Safaryan, Peter RichtárikICML 2021 · 70 citations
- Memory-Efficient LLM Pretraining via Minimalist Optimizer DesignAthanasios Glentis, Jiaxiang Li, Andi Han, Mingyi HongICML 2026 · 9 citations
