Neural ODE Transformers: Analyzing Internal Dynamics and Adaptive Fine-tuning
Anh Tong, Thanh Nguyen-Tang, Dongeun Lee, Duc Nguyen, Toan M. Tran, David Leo Wright Hall, Cheongwoong Kang, Jaesik Choi
摘要
Recent advancements in large language models (LLMs) based on transformer architectures have sparked significant interest in understanding their inner workings. In this paper, we introduce a novel approach to modeling transformer architectures using highly flexible non-autonomous neural ordinary differential equations (ODEs). Our proposed model parameterizes all weights of attention and feedforward blocks through neural networks, expressing these weights as functions of a continuous layer index. Through spectral analysis of the model's dynamics, we uncover an increase in eigenvalue magnitude that challenges the weightsharing assumption prevalent in existing theoretical studies. We also leverage the Lyapunov exponent to examine token-level sensitivity, enhancing model interpretability. Our neural ODE transformer demonstrates performance comparable to or better than vanilla transformers across various configurations and datasets, while offering flexible fine-tuning capabilities that can adapt to different architectural constraints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Semantic Tube Prediction: Beating LLM Data Efficiency with JEPAHai Huang, Yann LeCun, Randall BalestrieroICML 2026 · 被引用 8 次
- Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from JacobiansAkiyoshi Tomihari, Ryo KarakidaNeurIPS 2025 · 被引用 5 次
- IIET: Efficient Numerical Transformer via Implicit Iterative Euler MethodXinyu Liu, Bei Li, Jiahao Liu, Junhao Ruan 等EMNLP 2025
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
相关 Paper
- Go with the flow: Adaptive control for Neural ODEsMathieu Chalvidal, Matthew Ricci, Rufin VanRullen, Thomas SerreICLR 2021 · 被引用 2 次
- A Solvable Attention for Neural Scaling LawsBochen Lyu, Di Wang, Zhanxing ZhuICLR 2025
- Stateful ODE-Nets using Basis Function ExpansionsAlejandro F. Queiruga, N. Benjamin Erichson, Liam Hodgkinson, Michael W. MahoneyNeurIPS 2021 · 被引用 18 次
- Learning to Encode Position for Transformer with Continuous Dynamical ModelXuanqing Liu, Hsiang-Fu Yu, Inderjit S. Dhillon, Cho-Jui HsiehICML 2020 · 被引用 139 次
- Transformer Block Coupling and its Correlation with Generalization in LLMsMurdock Aubry, Haoming Meng, Anton Sugolov, Vardan PapyanICLR 2025
