Leveraging Relaxed Equilibrium by Lazy Transition for Sequence Modeling
Xi Ai, Bin Fang
Abstract
In sequence modeling, certain tokens are usually less ambiguous than others, and representations of these tokens require fewer refinements for disambiguation. However, given the nature of attention-based models like Transformer and UT (universal transformer), all tokens are equally processed towards depth. Inspired by the equilibrium phenomenon, we present a lazy transition, a mechanism to adjust the significance of iterative refinements for each token representation. Our lazy transition is deployed on top of UT to build LT (lazy transformer), where all tokens are processed unequally towards depth. Eventually, LT is encouraged to oscillate around a relaxed equilibrium. Our experiments show that LT outperforms baseline models on several tasks of machine translation, pre-training, Learning to Execute, and LAMBADA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on5
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 264 citations
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer et al.ACL 2020 · 210 citations
- Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual TranslationBiao Zhang, Ankur Bapna, Rico Sennrich, Orhan FiratICLR 2021 · 97 citations
Related papers
- Token-level Adaptive Training for Neural Machine TranslationShuhao Gu, Jinchao Zhang, Fandong Meng, Yang Feng et al.EMNLP 2020 · 32 citations
- Token Dropping for Efficient BERT PretrainingLe Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu et al.ACL 2022
- LightToken: A Task and Model-agnostic Lightweight Token Embedding Framework for Pre-trained Language ModelsHaoyu Wang, Ruirui Li, Haoming Jiang, Zhengyang Wang et al.KDD 2023 · 5 citations
- Looking Beyond the Top-1: Transformers Determine Top Tokens in OrderDaria Lioubashevski, Tomer Schlank, Gabriel Stanovsky, Ariel GoldsteinICML 2025
- Deep Transformers with Latent DepthXian Li, Asa Cooper Stickland, Yuqing Tang, Xiang KongNeurIPS 2020 · 32 citations
