RETVec: Resilient and Efficient Text Vectorizer
Elie Bursztein, Marina Zhang, Owen Vallis, Xinyu Jia, Alexey Kurakin
Abstract
This paper describes RETVec, an efficient, resilient, and multilingual text vectorizer designed for neural-based text processing. RETVec combines a novel character encoding with an optional small embedding model to embed words into a 256-dimensional vector space. The RETVec embedding model is pre-trained using pair-wise metric learning to be robust against typos and character-level adversarial attacks. In this paper, we evaluate and compare RETVec to state-of-the-art vectorizers and word embeddings on popular model architectures and datasets. These comparisons demonstrate that RETVec leads to competitive, multilingual models that are significantly more resilient to typos and adversarial text attacks. RETVec is available under the Apache 2 license at https://github.com/google-research/retvec.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98b729c2-4ae1-42e9-a4b4-459146e0214fCited by top-tier papers2
- Enhancing Large Language Models through Adaptive TokenizersMengyu Zheng, Hanting Chen, Tianyu Guo, Chong Zhu et al.NeurIPS 2024 · 11 citations
- RETSim: Resilient and Efficient Text SimilarityMarina Zhang, Owen S. Vallis, Aysegul Bumin, Tanay Vakharia et al.ICLR 2024 · 3 citations
Builds on3
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse LanguagesJack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie et al.ACL 2023 · 88 citations
- Circle Loss: A Unified Perspective of Pair Similarity OptimizationYifan Sun, Changmao Cheng, Yuhan Zhang, Chi Zhang et al.CVPR 2020
Related papers
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 200 citations
- Language-Agnostic Visual-Semantic EmbeddingsJonatas Wehrmann, Maurício Armani Lopes, Douglas M. Souza, Rodrigo C. BarrosICCV 2019 · 56 citations
- Joint Character-Level Word Embedding and Adversarial Stability Training to Defend Adversarial TextHui Liu, Yongzheng Zhang, Yipeng Wang, Zheng Lin et al.AAAI 2020 · 43 citations
- Searching for an Effective Defender: Benchmarking Defense against Adversarial Word SubstitutionZongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li et al.EMNLP 2021 · 46 citations
- Robust Encodings: A Framework for Combating Adversarial TyposErik Jones, Robin Jia, Aditi Raghunathan, Percy LiangACL 2020 · 92 citations
