Bridging Subword Gaps in Pretrain-Finetune Paradigm for Natural Language Generation
Xin Liu, Baosong Yang, Dayiheng Liu, Haibo Zhang, Weihua Luo, Min Zhang, Haiying Zhang, Jinsong Su
Abstract
A well-known limitation in pretrain-finetune paradigm lies in its inflexibility caused by the one-size-fits-all vocabulary. This potentially weakens the effect when applying pretrained models into natural language generation (NLG) tasks, especially for the subword distributions between upstream and downstream tasks with significant discrepancy. Towards approaching this problem, we extend the vanilla pretrain-finetune pipeline with an extra embedding transfer step. Specifically, a plug-and-play embedding generator is introduced to produce the representation of any input token, according to pre-trained embeddings of its morphologically similar ones. Thus, embeddings of mismatch tokens in downstream tasks can also be efficiently initialized. We conduct experiments on a variety of NLG tasks under the pretrain-finetune fashion. Experimental results and extensive analyses show that the proposed strategy offers us opportunities to feel free to transfer the vocabulary, leading to more efficient and better performed downstream NLG models. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72e48843-254f-453d-b248-4db0d8c4a8d2Cited by top-tier papers6
- Frequency-Aware Contrastive Learning for Neural Machine TranslationTong Zhang, Wei Ye, Baosong Yang, Long Zhang et al.AAAI 2022 · 35 citations
- ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine TranslationZhaocong Li, Xuebo Liu, Derek F. Wong, Lidia S. Chao et al.EMNLP 2022 · 20 citations
- Task-Adaptive Tokenization: Enhancing Long-Form Text Generation Efficacy in Mental Health and BeyondSiyang Liu, Naihao Deng, Sahand Sabour, Yilin Jia et al.EMNLP 2023 · 7 citations
- A Learning Rate Path Switching Training Paradigm for Version Updates of Large Language ModelsZhihao Wang, Shiyu Liu, Jianheng Huang, Wang Zheng et al.EMNLP 2024
- Mitigating Structural Knowledge Collapse in Domain-Specific LLMs via Morpheme-Aware KV-AggregationYuxuan Si, Zheqi Lv, Chengxi Zang, Zhengyu Chen et al.ACL 2026
Builds on7
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Incorporating BERT into Neural Machine TranslationJinhua Zhu, Yingce Xia, Lijun Wu, Di He et al.ICLR 2020 · 391 citations
- On the Inference Calibration of Neural Machine TranslationShuo Wang, Zhaopeng Tu, Shuming Shi, Yang LiuACL 2020 · 66 citations
- In Neural Machine Translation, What Does Transfer Learning Transfer?Alham Fikri Aji, Nikolay Bogoychev, Kenneth Heafield, Rico SennrichACL 2020 · 56 citations
- Improving Tree-Structured Decoder Training for Code Generation via Mutual LearningBinbin Xie, Jinsong Su, Yubin Ge, Xiang Li et al.AAAI 2021 · 30 citations
Related papers
- Plug-and-Play Knowledge Injection for Pre-trained Language ModelsZhengyan Zhang, Zhiyuan Zeng, Yankai Lin, Huadong Wang et al.ACL 2023 · 10 citations
- Different Strokes for Different Folks: Investigating Appropriate Further Pre-training Approaches for Diverse Dialogue TasksYao Qiu, Jinchao Zhang, Jie ZhouEMNLP 2021 · 1 citation
- How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language ModelsPhillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder et al.ACL 2021
- Scaling Laws for Forgetting during Finetuning with Pretraining Data InjectionLouis Béthune, David Grangier, Dan Busbridge, Eleonora Gualdoni et al.ICML 2025
- LangBridge: Interpreting Image as a Combination of Language EmbeddingsJiaqi Liao, Yuwei Niu, Fanqing Meng, Hao Li et al.ICCV 2025
