Elevating Code-mixed Text Handling through Auditory Information of Words
Mamta, Zishan Ahmad, Asif Ekbal
摘要
With the growing popularity of code-mixed data, there is an increasing need for better handling of this type of data, which poses a number of challenges, such as dealing with spelling variations, multiple languages, different scripts, and a lack of resources. Current language models face difficulty in effectively handling code-mixed data as they primarily focus on the semantic representation of words and ignore the auditory phonetic features. This leads to difficulties in handling spelling variations in code-mixed text. In this paper, we propose an effective approach for creating language models for handling code-mixed textual data using auditory information of words from SOUNDEX. Our approach includes a pre-training step based on masked-language-modelling, which includes SOUNDEX representations (SAMLM) and a new method of providing input data to the pretrained model. Through experimentation on various code-mixed datasets (of different languages) for sentiment, offensive and aggression classification tasks, we establish that our novel language modeling approach (SAMLM) results in improved robustness towards adversarial attacks on code-mixed classification tasks. Additionally, our SAMLM based approach also results in better classification results over the popular baselines for code-mixed tasks. We use the explainability technique, SHAP (SHapley Additive exPlanations) to explain how the auditory features incorporated through SAMLM assist the model to handle the code-mixed text effectively and increase robustness against adversarial attacks 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- BiasWipe: Mitigating Unintended Bias in Text Classifiers through Model InterpretabilityMamta Mamta, Rishikant Chigrupaatii, Asif EkbalEMNLP 2024 · 被引用 4 次
- Explainability and Interpretability of Multilingual Large Language Models: A SurveyLucas Resck, Isabelle Augenstein, Anna KorhonenEMNLP 2025
它引用的顶会 Paper2
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
相关 Paper
- Improving Pretraining Techniques for Code-Switched NLPRicheek Das, Sahasra Ranjan, Shreya Pathak, Preethi JyothiACL 2023 · 被引用 3 次
- SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMsBoyi Deng, Yu Wan, Baosong Yang, Fei Huang 等ICLR 2026 · 被引用 2 次
- GenEx: A Commonsense-aware Unified Generative Framework for Explainable Cyberbullying DetectionKrishanu Maity, Raghav Jain, Prince Jha, Sriparna Saha 等EMNLP 2023 · 被引用 4 次
- Learning from Multiple Noisy Augmented Data Sets for Better Cross-Lingual Spoken Language UnderstandingYingmei Guo, Linjun Shou, Jian Pei, Ming Gong 等EMNLP 2021 · 被引用 2 次
- SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic SoundscapesTony Alex, Sara Atito, Armin Mustafa, Muhammad Awais 等ICLR 2025
