IOT: Instance-wise Layer Reordering for Transformer Structures
Jinhua Zhu, Lijun Wu, Yingce Xia, Shufang Xie, Tao Qin, Wengang Zhou, Houqiang Li, Tie-Yan Liu
Abstract
With sequentially stacked self-attention, (optional) encoder-decoder attention, and feed-forward layers, Transformer achieves big success in natural language processing (NLP), and many variants have been proposed. Currently, almost all these models assume that the layer order is fixed and kept the same across data samples. We observe that different data samples actually favor different orders of the layers. Based on this observation, in this work, we break the assumption of the fixed layer order in the Transformer and introduce instance-wise layer reordering into the model structure. Our Instance-wise Ordered Transformer (IOT) can model variant functions by reordered layers, which enables each sample to select the better one to improve the model performance under the constraint of almost the same number of parameters. To achieve this, we introduce a light predictor with negligible parameter and inference cost to decide the most capable and favorable layer order for any input sequence. Experiments on 3 tasks (neural machine translation, abstractive summarization, and code generation) and 9 datasets demonstrate consistent improvements of our method. We further show that our method can also be applied to other architectures beyond Transformer. Our code is released at Github.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- Depth-Adaptive TransformerMaha Elbayad, Jiatao Gu, Edouard Grave, Michael AuliICLR 2020 · 264 citations
- FastBERT: a Self-distilling BERT with Adaptive Inference TimeWeijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao et al.ACL 2020 · 257 citations
- Graph Transformer for Graph-to-Sequence LearningDeng Cai, Wai LamAAAI 2020 · 247 citations
Related papers
- Improving Transformer Models by Reordering their SublayersOfir Press, Noah A. Smith, Omer LevyACL 2020 · 6 citations
- AMOM: Adaptive Masking over Masking for Conditional Masked Language ModelYisheng Xiao, Ruiyang Xu, Lijun Wu, Juntao Li et al.AAAI 2023 · 14 citations
- Guiding Non-Autoregressive Neural Machine Translation Decoding with Reordering InformationQiu Ran, Yankai Lin, Peng Li, Jie ZhouAAAI 2021 · 82 citations
- Transformer Layers as PaintersQi Sun, Marc Pickett, Aakash Kumar Nain, Llion JonesAAAI 2025 · 49 citations
- Sequence Generation with Mixed RepresentationsLijun Wu, Shufang Xie, Yingce Xia, Yang Fan et al.ICML 2020 · 18 citations
