ODE Transformer: An Ordinary Differential Equation-Inspired Model for Sequence Generation
Bei Li, Quan Du, Tao Zhou, Yi Jing, Shuhan Zhou, Xin Zeng, Tong Xiao, Jingbo Zhu, Xuebo Liu, Min Zhang
Abstract
Residual networks are an Euler discretization of solutions to Ordinary Differential Equations (ODE). This paper explores a deeper relationship between Transformer and numerical ODE methods. We first show that a residual block of layers in Transformer can be described as a higher-order solution to ODE. Inspired by this, we design a new architecture, ODE Transformer, which is analogous to the Runge-Kutta method that is well motivated in ODE. As a natural extension to Transformer, ODE Transformer is easy to implement and efficient to use. Experimental results on the large-scale machine translation, abstractive summarization, and grammar error correction tasks demonstrate the high genericity of ODE Transformer. It can gain large improvements in model performance over strong baselines (e.g., 30.77 and 44.11 BLEU scores on the WMT’14 English-German and English-French benchmarks) at a slight cost in inference efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 534f4652-7569-4b04-a9cc-d418ea4b88beCited by top-tier papers11
- TemplateGEC: Improving Grammatical Error Correction with Detection TemplateYinghao Li, Xuebo Liu, Shuo Wang, Peiyuan Gong et al.ACL 2023 · 21 citations
- Learning Multiscale Transformer Models for Sequence GenerationBei Li, Tong Zheng, Yi Jing, Chengbo Jiao et al.ICML 2022 · 15 citations
- Revisiting Grammatical Error Correction Evaluation and BeyondPeiyuan Gong, Xuebo Liu, Heyan Huang, Min ZhangEMNLP 2022 · 11 citations
- Optimal Control for Transformer Architectures: Enhancing Generalization, Robustness and EfficiencyKelvin Kan, Xingjian Li, Benjamin J. Zhang, Tuhin Sahai et al.NeurIPS 2025 · 9 citations
- Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient LearningBei Li, Tong Zheng, Rui Wang, Jiahao Liu et al.NeurIPS 2024 · 5 citations
Builds on13
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Incorporating BERT into Neural Machine TranslationJinhua Zhu, Yingce Xia, Lijun Wu, Di He et al.ICLR 2020 · 391 citations
- Lite Transformer with Long-Short Range AttentionZhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin et al.ICLR 2020 · 379 citations
- Improving Transformer Optimization Through Better InitializationXiao Shi Huang, Felipe Pérez, Jimmy Ba, Maksims VolkovsICML 2020 · 181 citations
Related papers
- Learning to Encode Position for Transformer with Continuous Dynamical ModelXuanqing Liu, Hsiang-Fu Yu, Inderjit S. Dhillon, Cho-Jui HsiehICML 2020 · 139 citations
- IIET: Efficient Numerical Transformer via Implicit Iterative Euler MethodXinyu Liu, Bei Li, Jiahao Liu, Junhao Ruan et al.EMNLP 2025
- HNO: High-Order Numerical Architecture for ODE-Inspired Deep Unfolding NetworksLin Kong, Wei Sun, Fanhua Shang, Yuanyuan Liu et al.AAAI 2022 · 1 citation
- ResNet After All: Neural ODEs and Their Numerical SolutionKatharina Ott, Prateek Katiyar, Philipp Hennig, Michael TiemannICLR 2021 · 34 citations
- Do Residual Neural Networks discretize Neural Ordinary Differential Equations?Michael E. Sander, Pierre Ablin, Gabriel PeyréNeurIPS 2022 · 42 citations
