LEA: Improving Sentence Similarity Robustness to Typos Using Lexical Attention Bias
Mario Almagro, Emilio J. Almazán, Diego Ortego, David Jiménez
Abstract
Textual noise, such as typos or abbreviations, is a well-known issue that penalizes vanilla Transformers for most downstream tasks. We show that this is also the case for sentence similarity, a fundamental task in multiple domains, e.g. matching, retrieval or paraphrasing. Sentence similarity can be approached using cross-encoders, where the two sentences are concatenated in the input allowing the model to exploit the inter-relations between them. Previous works addressing the noise issue mainly rely on data augmentation strategies, showing improved robustness when dealing with corrupted samples that are similar to the ones used for training. However, all these methods still suffer from the token distribution shift induced by typos. In this work, we propose to tackle textual noise by equipping cross-encoders with a novel LExical-aware Attention module (LEA) that incorporates lexical similarities between words in both sentences. By using raw text similarities, our approach avoids the tokenization shift problem obtaining improved robustness. We demonstrate that the attention bias introduced by LEA helps cross-encoders to tackle complex scenarios with textual noise, specially in domains with short-text descriptions and limited context. Experiments using three popular Transformer encoders in five e-commerce datasets for product matching show that LEA consistently boosts performance under the presence of noise, while remaining competitive on the original (clean) splits. We also evaluate our approach in two datasets for textual entailment and paraphrasing showing that LEA is robust to typos in domains with longer sentences and more natural context. Additionally, we thoroughly analyze several design choices in our approach, providing insights about the impact of the decisions made and fostering future research in cross-encoders dealing with typos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- TacoPrompt: A Collaborative Multi-Task Prompt Learning Method for Self-Supervised Taxonomy CompletionHongyuan Xu, Ciyi Liu, Yuhang Niu, Yunong Chen et al.EMNLP 2023 · 9 citations
- Investigating Neurons and Heads in Transformer-based LLMs for Typographical ErrorsKohei Tsuji, Tatsuya Hiraoka, Yuchang Cheng, Eiji Aramaki et al.EMNLP 2025
Builds on16
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
Related papers
- Error Norm Truncation: Robust Training in the Presence of Data Noise for Text Generation ModelsTianjian Li, Haoran Xu, Philipp Koehn, Daniel Khashabi et al.ICLR 2024 · 6 citations
- NAT: Noise-Aware Training for Robust Neural Sequence LabelingMarcin Namysl, Sven Behnke, Joachim KöhlerACL 2020
- DenoSent: A Denoising Objective for Self-Supervised Sentence Representation LearningXinghao Wang, Junliang He, Pengyu Wang, Yunhua Zhou et al.AAAI 2024 · 11 citations
- OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence EmbeddingZhan Shi, Guoyin Wang, Ke Bai, Jiwei Li et al.EMNLP 2023 · 3 citations
- Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language LearningShivaen Ramshetty, Gaurav Verma, Srijan KumarACL 2023 · 4 citations
