Local Byte Fusion for Neural Machine Translation
Makesh Narsimhan Sreedhar, Xiangpeng Wan, Yu Cheng, Junjie Hu
摘要
Subword tokenization schemes are the dominant technique used in current NLP models. However, such schemes can be rigid and tokenizers built on one corpus may not adapt well to other parallel corpora. It has also been observed that in multilingual corpora, subword tokenization schemes oversegment low-resource languages, leading to a drop in translation performance. An alternative to subword tokenizers is byte-based tokenization, i.e., tokenization into byte sequences using the UTF-8 encoding scheme. Byte tokens often represent inputs at a sub-character granularity, i.e., one character can be represented by a span of byte tokens. This results in much longer byte sequences that are hard to interpret without aggregating local information from multiple byte tokens. In this paper, we propose a Local Byte Fusion (LOBEF) method for byte-based machine translation—utilizing byte n-gram and word boundaries—to aggregate local semantic information. Extensive experiments on multilingual translation, zero-shot cross-lingual transfer, and domain adaptation reveal a consistent improvement over vanilla byte-based models. Further analysis also indicates that our byte-based models are parameter-efficient and perform competitive to subword models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Dynamic Chunking for End-to-End Hierarchical Sequence ModelingSukjun Hwang, Brandon Wang, Albert GuICLR 2026 · 被引用 76 次
- SpaceByte: Towards Deleting Tokenization from Large Language ModelingKevin SlagleNeurIPS 2024 · 被引用 34 次
- Enhancing Large Language Models through Adaptive TokenizersMengyu Zheng, Hanting Chen, Tianyu Guo, Chong Zhu 等NeurIPS 2024 · 被引用 11 次
- DNACHUNKER: Learnable Tokenization for DNA Language ModelsTaewon Kim, Jihwan Shin, Hyomin Kim, Youngmok Jung 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper5
- Neural Machine Translation with Byte-Level SubwordsChanghan Wang, Kyunghyun Cho, Jiatao GuAAAI 2020 · 被引用 213 次
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta 等ICLR 2022 · 被引用 198 次
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 被引用 17 次
- Unsupervised Domain Clusters in Pretrained Language ModelsRoee Aharoni, Yoav GoldbergACL 2020 · 被引用 13 次
- How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language ModelsPhillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder 等ACL 2021
相关 Paper
- MAGNET: Improving the Multilingual Fairness of Language Models with Adaptive Gradient-Based TokenizationOrevaoghene Ahia, Sachin Kumar, Hila Gonen, Valentin Hofmann 等NeurIPS 2024 · 被引用 37 次
- ByteFlow: Language Modeling through Adaptive Byte Compression without a TokenizerChunyuan Deng, Sanket Lokegaonkar, Colin Lockard, Besnik Fetahu 等ICLR 2026 · 被引用 1 次
- Beyond Perplexity: UTF-8 Validity in Byte-aware Language ModelsSangwhan Moon, Daisuke Oba, Youmi Ma, Tatsuya Hiraoka 等ICML 2026
- Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language ModelsPit Neitemeier, Björn Deiseroth, Constantin Eichenberg, Lukas BallesICLR 2025
- Corpus-Dependent Subcharacter Encoding via HMM-Guided Code AssignmentTatsuya HiraokaACL 2026
