Lune

EMNLP2021Top-tier venue

Robust Open-Vocabulary Translation from Visual Text Representations

Elizabeth Salesky, David Etter, Matt Post

2021Year
33Citations
11Top-tier citations

Abstract

Machine translation models have discrete vo cabularies and commonly use subword seg mentation techniques to achieve an 'open vo cabulary.' This approach relies on consis tent and correct underlying unicode sequences, and makes models susceptible to degrada tion from common types of noise and vari ation. Motivated by the robustness of hu man language processing, we propose the use of visual text representations, which dispense with a finite set of text embeddings in favor of continuous vocabularies created by process ing visually rendered text with sliding win dows. We show that models using visual text representations approach or match per formance of traditional text models on small and larger datasets. More importantly, mod els with visual embeddings demonstrate sig nificant robustness to varied types of noise, achieving e.g., 25.9 BLEU on a character per muted German-English task where subword models degrade to 1.9.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 66d4b910-d033-490d-83c4-9fb8b3acd648

Cited by top-tier papers11

Ask how each one uses it

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines