Reducing Per-Sample Harm in Stochastic Optimization
Apostolos Avranas
Abstract
Modern optimizers combine gradients from the current mini-batch with historical optimization state, such as momentum or adaptive moments. While highly effective, aggregating across the batch and incorporating this history can produce parameter updates that increase the loss of individual samples. We term this effect harm and formalize the parameter update as an optimization problem that explicitly minimizes the conflicting impact of both batch averaging and past optimization state on current data. Because the exact formulation is intractable, we introduce a highly efficient proxy. We first reduce the problem's dimensionality to the batch size, and then drastically cut memory and speed bottlenecks by successfully restricting the optimization to the last linear layer. This hinges on the unexpected finding that this layer alone reliably captures the second-order statistics of the per-sample gradients . The resulting surrogate problem integrates readily into standard optimizers like SGD and AdamW, and can be solved using a small number of GPU-friendly iterations. Crucially, the method exhibits favorable scaling properties , as the relative computational overhead shrinks as the model size or input grows. Experiments on image classification benchmarks confirm reduced per-sample interference and improved generalization .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a12ef05-b195-4ed3-902a-a68eaa6e2d3dBuilds on18
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
Related papers
- The AdEMAMix Optimizer: Better, Faster, OlderMatteo Pagliardini, Pierre Ablin, David GrangierICLR 2025
- Friendly Sharpness-Aware MinimizationTao Li, Pan Zhou, Zhengbao He, Xinwen Cheng et al.CVPR 2024
- Memory-Efficient LLM Pretraining via Minimalist Optimizer DesignAthanasios Glentis, Jiaxiang Li, Andi Han, Mingyi HongICML 2026 · 9 citations
- Adam Reduces a Unique Form of Sharpness: Theoretical Insights Near the Minimizer ManifoldXinghan Li, Haodong Wen, Kaifeng LyuNeurIPS 2025 · 6 citations
- IO-Adam: Rethinking Memory-Efficient Adaptive Optimizers from Gradient ComputationYiting Chen, Zongwei Huo, Junchi YanICML 2026
