Bridging Subword Gaps in Pretrain-Finetune Paradigm for Natural Language Generation
Xin Liu, Baosong Yang, Dayiheng Liu, Haibo Zhang, Weihua Luo, Min Zhang, Haiying Zhang, Jinsong Su
摘要
A well-known limitation in pretrain-finetune paradigm lies in its inflexibility caused by the one-size-fits-all vocabulary. This potentially weakens the effect when applying pretrained models into natural language generation (NLG) tasks, especially for the subword distributions between upstream and downstream tasks with significant discrepancy. Towards approaching this problem, we extend the vanilla pretrain-finetune pipeline with an extra embedding transfer step. Specifically, a plug-and-play embedding generator is introduced to produce the representation of any input token, according to pre-trained embeddings of its morphologically similar ones. Thus, embeddings of mismatch tokens in downstream tasks can also be efficiently initialized. We conduct experiments on a variety of NLG tasks under the pretrain-finetune fashion. Experimental results and extensive analyses show that the proposed strategy offers us opportunities to feel free to transfer the vocabulary, leading to more efficient and better performed downstream NLG models. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Frequency-Aware Contrastive Learning for Neural Machine TranslationTong Zhang, Wei Ye, Baosong Yang, Long Zhang 等AAAI 2022 · 被引用 35 次
- ConsistTL: Modeling Consistency in Transfer Learning for Low-Resource Neural Machine TranslationZhaocong Li, Xuebo Liu, Derek F. Wong, Lidia S. Chao 等EMNLP 2022 · 被引用 20 次
- Task-Adaptive Tokenization: Enhancing Long-Form Text Generation Efficacy in Mental Health and BeyondSiyang Liu, Naihao Deng, Sahand Sabour, Yilin Jia 等EMNLP 2023 · 被引用 7 次
- A Learning Rate Path Switching Training Paradigm for Version Updates of Large Language ModelsZhihao Wang, Shiyu Liu, Jianheng Huang, Wang Zheng 等EMNLP 2024
- Mitigating Structural Knowledge Collapse in Domain-Specific LLMs via Morpheme-Aware KV-AggregationYuxuan Si, Zheqi Lv, Chengxi Zang, Zhengyu Chen 等ACL 2026
它引用的顶会 Paper7
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Incorporating BERT into Neural Machine TranslationJinhua Zhu, Yingce Xia, Lijun Wu, Di He 等ICLR 2020 · 被引用 391 次
- On the Inference Calibration of Neural Machine TranslationShuo Wang, Zhaopeng Tu, Shuming Shi, Yang LiuACL 2020 · 被引用 66 次
- In Neural Machine Translation, What Does Transfer Learning Transfer?Alham Fikri Aji, Nikolay Bogoychev, Kenneth Heafield, Rico SennrichACL 2020 · 被引用 56 次
- Improving Tree-Structured Decoder Training for Code Generation via Mutual LearningBinbin Xie, Jinsong Su, Yubin Ge, Xiang Li 等AAAI 2021 · 被引用 30 次
相关 Paper
- Plug-and-Play Knowledge Injection for Pre-trained Language ModelsZhengyan Zhang, Zhiyuan Zeng, Yankai Lin, Huadong Wang 等ACL 2023 · 被引用 10 次
- Different Strokes for Different Folks: Investigating Appropriate Further Pre-training Approaches for Diverse Dialogue TasksYao Qiu, Jinchao Zhang, Jie ZhouEMNLP 2021 · 被引用 1 次
- How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language ModelsPhillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder 等ACL 2021
- Scaling Laws for Forgetting during Finetuning with Pretraining Data InjectionLouis Béthune, David Grangier, Dan Busbridge, Eleonora Gualdoni 等ICML 2025
- LangBridge: Interpreting Image as a Combination of Language EmbeddingsJiaqi Liao, Yuwei Niu, Fanqing Meng, Hao Li 等ICCV 2025
