Frequency-aware SGD for Efficient Embedding Learning with Provable Benefits
Yan Li, Dhruv Choudhary, Xiaohan Wei, Baichuan Yuan, Bhargav Bhushanam, Tuo Zhao, Guanghui Lan
Abstract
Embedding learning has found widespread applications in recommendation systems and natural language modeling, among other domains. To learn quality embeddings efficiently, adaptive learning rate algorithms have demonstrated superior empirical performance over SGD, largely accredited to their token-dependent learning rate. However, the underlying mechanism for the efficiency of token-dependent learning rate remains underexplored. We show that incorporating frequency information of tokens in the embedding learning problems leads to provably efficient algorithms, and demonstrate that common adaptive algorithms implicitly exploit the frequency information to a large extent. Specifically, we propose (Counterbased) Frequency-aware Stochastic Gradient Descent, which applies a frequency-dependent learning rate for each token, and exhibits provable speed-up compared to SGD when the token distribution is imbalanced. Empirically, we show the proposed algorithms are able to improve or match adaptive algorithms on benchmark recommendation tasks and a large-scale industrial recommendation system, closing the performance gap between SGD and adaptive algorithms, while using significantly lower memory. Our results are the first to show token-dependent learning rate provably improves convergence for non-convex embedding learning problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- AdaEmbed: Adaptive Embedding for Large-Scale Recommendation ModelsFan Lai, Wei Zhang, Rui Liu, William Tsai et al.OSDI 2023 · 23 citations
- Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's LawFrederik Kunstner, Francis BachNeurIPS 2025 · 21 citations
- AdaGrad under Anisotropic SmoothnessYuxing Liu, Rui Pan, Tong ZhangICLR 2025
Builds on2
Related papers
- Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language ModelsFrederik Kunstner, Alan Milligan, Robin Yadav, Mark Schmidt et al.NeurIPS 2024 · 100 citations
- Frequency-Aware Contrastive Learning for Neural Machine TranslationTong Zhang, Wei Ye, Baosong Yang, Long Zhang et al.AAAI 2022 · 35 citations
- FEC: Efficient Deep Recommendation Model Training with Flexible Embedding CommunicationKaihao Ma, Xiao Yan, Zhenkun Cai, Yuzhen Huang et al.SIGMOD 2023 · 8 citations
- Zipfian WhiteningSho Yokoi, Han Bao, Hiroto Kurita, Hidetoshi ShimodairaNeurIPS 2024 · 3 citations
- Rare Tokens Degenerate All Tokens: Improving Neural Text Generation via Adaptive Gradient Gating for Rare Token EmbeddingsSangwon Yu, Jongyoon Song, Heeseung Kim, Seongmin Lee et al.ACL 2022 · 42 citations
