Better Embeddings with Coupled Adam
Felix Stollenwerk, Tobias Stollenwerk
2025年份
4被引次数
1顶会引用
摘要
Despite their remarkable capabilities, LLMs learn word representations that exhibit the undesirable yet poorly understood feature of anisotropy. In this paper, we argue that the second moment in Adam is a cause of anisotropic embeddings, and suggest a modified optimizer called Coupled Adam to mitigate the problem. Our experiments demonstrate that Coupled Adam significantly improves the quality of embeddings, while also leading to better upstream and downstream performance on large enough datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Improving Neural Language Generation with Spectrum ControlLingxiao Wang, Jing Huang, Kevin Huang, Ziniu Hu 等ICLR 2020 · 被引用 94 次
- Rare Tokens Degenerate All Tokens: Improving Neural Text Generation via Adaptive Gradient Gating for Rare Token EmbeddingsSangwon Yu, Jongyoon Song, Heeseung Kim, Seongmin Lee 等ACL 2022 · 被引用 42 次
相关 Paper
- Stable Anisotropic RegularizationWilliam Rudman, Carsten EickhoffICLR 2024 · 被引用 13 次
- IsoBN: Fine-Tuning BERT with Isotropic Batch NormalizationWenxuan Zhou, Bill Yuchen Lin, Xiang RenAAAI 2021 · 被引用 29 次
- MKOR: Momentum-Enabled Kronecker-Factor-Based Optimizer Using Rank-1 UpdatesMohammad Mozaffari, Sikan Li, Zhao Zhang, Maryam Mehri DehnaviNeurIPS 2023 · 被引用 7 次
- AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-trainingHuishuai Zhang, Bohan Wang, Luoxin ChenEMNLP 2025 · 被引用 1 次
- From Attention to Activation: Unraveling the Enigmas of Large Language ModelsPrannay Kaul, Chengcheng Ma, Ismail Elezi, Jiankang DengICLR 2025
