RETVec: Resilient and Efficient Text Vectorizer
Elie Bursztein, Marina Zhang, Owen Vallis, Xinyu Jia, Alexey Kurakin
摘要
This paper describes RETVec, an efficient, resilient, and multilingual text vectorizer designed for neural-based text processing. RETVec combines a novel character encoding with an optional small embedding model to embed words into a 256-dimensional vector space. The RETVec embedding model is pre-trained using pair-wise metric learning to be robust against typos and character-level adversarial attacks. In this paper, we evaluate and compare RETVec to state-of-the-art vectorizers and word embeddings on popular model architectures and datasets. These comparisons demonstrate that RETVec leads to competitive, multilingual models that are significantly more resilient to typos and adversarial text attacks. RETVec is available under the Apache 2 license at https://github.com/google-research/retvec.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Enhancing Large Language Models through Adaptive TokenizersMengyu Zheng, Hanting Chen, Tianyu Guo, Chong Zhu 等NeurIPS 2024 · 被引用 11 次
- RETSim: Resilient and Efficient Text SimilarityMarina Zhang, Owen S. Vallis, Aysegul Bumin, Tanay Vakharia 等ICLR 2024 · 被引用 3 次
它引用的顶会 Paper3
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li 等NDSS 2019 · 被引用 876 次
- MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse LanguagesJack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie 等ACL 2023 · 被引用 88 次
- Circle Loss: A Unified Perspective of Pair Similarity OptimizationYifan Sun, Changmao Cheng, Yuhan Zhang, Chi Zhang 等CVPR 2020
相关 Paper
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 被引用 200 次
- Language-Agnostic Visual-Semantic EmbeddingsJonatas Wehrmann, Maurício Armani Lopes, Douglas M. Souza, Rodrigo C. BarrosICCV 2019 · 被引用 56 次
- Joint Character-Level Word Embedding and Adversarial Stability Training to Defend Adversarial TextHui Liu, Yongzheng Zhang, Yipeng Wang, Zheng Lin 等AAAI 2020 · 被引用 43 次
- Searching for an Effective Defender: Benchmarking Defense against Adversarial Word SubstitutionZongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li 等EMNLP 2021 · 被引用 46 次
- Robust Encodings: A Framework for Combating Adversarial TyposErik Jones, Robin Jia, Aditi Raghunathan, Percy LiangACL 2020 · 被引用 92 次
