Sequence Generation with Mixed Representations
Lijun Wu, Shufang Xie, Yingce Xia, Yang Fan, Jian-Huang Lai, Tao Qin, Tie-Yan Liu
Abstract
Tokenization is the first step of many natural language processing (NLP) tasks and plays an important role for neural NLP models. Tokenization methods such as byte-pair encoding and Senten-cePiece, which can greatly reduce the large vocabulary size and deal with out-of-vocabulary words, have shown to be effective and are widely adopted for sequence generation tasks. While various tokenization methods exist, there is no common acknowledgement which one is the best. In this work, we propose to leverage the mixed representations from different tokenizers for sequence generation tasks, which can take the advantages of each individual tokenization method. Specifically, we introduce a new model architecture to incorporate mixed representations and a co-teaching algorithm to better utilize the diversity of different tokenization methods. Our approach achieves significant improvements on neural machine translation tasks with six language pairs, as well as an abstractive summarization task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f0fd9f6-b981-44a1-8a5c-a9dfd576b9bfCited by top-tier papers4
- BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine TranslationHaoran Xu, Benjamin Van Durme, Kenton W. MurrayEMNLP 2021 · 55 citations
- Learning Multiscale Transformer Models for Sequence GenerationBei Li, Tong Zheng, Yi Jing, Chengbo Jiao et al.ICML 2022 · 15 citations
- EM-Network: Oracle Guided Self-distillation for Sequence LearningJi Won Yoon, Sunghwan Ahn, Hyeonseung Lee, Minchan Kim et al.ICML 2023 · 3 citations
- CipherDAug: Ciphertext based Data Augmentation for Neural Machine TranslationNishant Kambhatla, Logan Born, Anoop SarkarACL 2022
Builds on1
Related papers
- Tokenization Is More Than CompressionCraig W. Schmidt, Varshini Reddy, Haoran Zhang, Alec Alameddine et al.EMNLP 2024 · 16 citations
- Twist Decoding: Diverse Generators Guide Each OtherJungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Hao Peng et al.EMNLP 2022
- Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in TokenizationNegar Foroutan, Clara Meister, Debjit Paul, Joel Niklaus et al.ACL 2026 · 14 citations
- A Partition Cover Approach to TokenizationJia Peng Lim, Shawn Tan, Davin Choo, Hady W. LauwNeurIPS 2025 · 6 citations
- Getting the most out of your tokenizer for pre-training and domain adaptationGautier Dagan, Gabriel Synnaeve, Baptiste RozièreICML 2024 · 68 citations
