Robust Open-Vocabulary Translation from Visual Text Representations
Elizabeth Salesky, David Etter, Matt Post
Abstract
Machine translation models have discrete vo cabularies and commonly use subword seg mentation techniques to achieve an 'open vo cabulary.' This approach relies on consis tent and correct underlying unicode sequences, and makes models susceptible to degrada tion from common types of noise and vari ation. Motivated by the robustness of hu man language processing, we propose the use of visual text representations, which dispense with a finite set of text embeddings in favor of continuous vocabularies created by process ing visually rendered text with sliding win dows. We show that models using visual text representations approach or match per formance of traditional text models on small and larger datasets. More importantly, mod els with visual embeddings demonstrate sig nificant robustness to varied types of noise, achieving e.g., 25.9 BLEU on a character per muted German-English task where subword models degrade to 1.9.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66d4b910-d033-490d-83c4-9fb8b3acd648Cited by top-tier papers11
- Vision-centric Token Compression in Large Language ModelLing Xing, Alex Jinpeng Wang, Rui Yan, Xiangbo Shu et al.NeurIPS 2025 · 32 citations
- Language Modelling with PixelsPhillip Rust, Jonas F. Lotz, Emanuele Bugliarello, Elizabeth Salesky et al.ICLR 2023 · 17 citations
- Did Translation Models Get More Robust Without Anyone Even Noticing?Ben Peters, André F. T. MartinsACL 2025 · 10 citations
- Proxy Compression for Language ModelingLin Zheng, Li Xinyu, Qian Liu, Xiachong Feng et al.ICML 2026 · 3 citations
- Text Rendering Strategies for Pixel Language ModelsJonas F. Lotz, Elizabeth Salesky, Phillip Rust, Desmond ElliottEMNLP 2023 · 3 citations
Builds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Neural Machine Translation with Byte-Level SubwordsChanghan Wang, Kyunghyun Cho, Jiatao GuAAAI 2020 · 213 citations
- BPE-Dropout: Simple and Effective Subword RegularizationIvan Provilkov, Dmitrii Emelianenko, Elena VoitaACL 2020 · 17 citations
- COMET: A Neural Framework for MT EvaluationRicardo Rei, Craig Stewart, Ana C. Farinha, Alon LavieEMNLP 2020 · 6 citations
- OCR Post Correction for Endangered Language TextsShruti Rijhwani, Antonios Anastasopoulos, Graham NeubigEMNLP 2020 · 1 citation
Related papers
- Globetrotter: Connecting Languages by Connecting ImagesDídac Surís, Dave Epstein, Carl VondrickCVPR 2022 · 7 citations
- From Characters to Words: Hierarchical Pre-trained Language Model for Open-vocabulary Language UnderstandingLi Sun, Florian Luisier, Kayhan Batmanghelich, Dinei A. F. Florêncio et al.ACL 2023
- Neural Machine Translation with Universal Visual RepresentationZhuosheng Zhang, Kehai Chen, Rui Wang, Masao Utiyama et al.ICLR 2020 · 117 citations
- Unsupervised Multimodal Neural Machine Translation with Pseudo Visual PivotingPo-Yao Huang, Junjie Hu, Xiaojun Chang, Alexander G. HauptmannACL 2020 · 43 citations
- Semantic Robustness Certification for Vision-Language ModelsPeiyu Yang, Paul MONTAGUE, Feng Liu, Andrew C. Cullen et al.ICML 2026
