Bi-Phone: Modeling Inter Language Phonetic Influences in Text
Abhirut Gupta, Ananya B. Sai, Richard Sproat, Yuri Vasilevski, James S. Ren, Ambarish Jash, Sukhdeep S. Sodhi, Aravindan Raghuveer
摘要
A large number of people are forced to use the Web in a language they have low literacy in due to technology asymmetries. Written text in the second language (L2) from such users often contains a large number of errors that are influenced by their native language (L1). We propose a method to mine phoneme confusions (sounds in L2 that an L1 speaker is likely to conflate) for pairs of L1 and L2. These confusions are then plugged into a generative model (Bi-Phone) for synthetically producing corrupted L2 text. Through human evaluations, we show that Bi-Phone generates plausible corruptions that differ across L1s and also have widespread coverage on the Web. We also corrupt the popular language understanding benchmark SuperGLUE with our technique (FunGLUE for Phonetically Noised GLUE) and show that SoTA language understating models perform poorly. We also introduce a new phoneme prediction pre-training task which helps byte models to recover performance close to SuperGLUE. Finally, we also release the FunGLUE benchmark to promote further research in phonetically robust language models. To the best of our knowledge, FunGLUE is the first benchmark to introduce L1-L2 interactions in text.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- XGLUE: A New Benchmark Datasetfor Cross-lingual Pre-training, Understanding and GenerationYaobo Liang, Nan Duan, Yeyun Gong, Ning Wu 等EMNLP 2020 · 被引用 232 次
- Sinhala Encoder-only Language Models and EvaluationTharindu Ranasinghe, Hansi Hettiarachchi, Nadeesha Chathurangi Naradde Vidana Pathirana, Damith Premasiri 等ACL 2025 · 被引用 5 次
- Superlim: A Swedish Language Understanding Evaluation BenchmarkAleksandrs Berdicevskis, Gerlof Bouma, Robin Kurtz, Felix Morger 等EMNLP 2023 · 被引用 2 次
- VALUE: Understanding Dialect Disparity in NLUCaleb Ziems, Jiaao Chen, Camille Harris, Jessica Anderson 等ACL 2022 · 被引用 57 次
- TEMA: Token Embeddings Mapping for Enriching Low-Resource Language ModelsRodolfo Zevallos, Núria Bel, Mireia FarrúsEMNLP 2024
