Speaker Information Can Guide Models to Better Inductive Biases: A Case Study On Predicting Code-Switching
Alissa Ostapenko, Shuly Wintner, Melinda Fricke, Yulia Tsvetkov
摘要
Natural language processing (NLP) models trained on people-generated data can be unreliable because, without any constraints, they can learn from spurious correlations that are not relevant to the task. We hypothesize that enriching models with speaker information in a controlled, educated way can guide them to pick up on relevant inductive biases. For the speaker-driven task of predicting code-switching points in English–Spanish bilingual dialogues, we show that adding sociolinguistically-grounded speaker features as prepended prompts significantly improves accuracy. We find that by adding influential phrases to the input, speaker-informed models learn useful and explainable linguistic information. To our knowledge, we are the first to incorporate speaker characteristics in a neural model for code-switching, and more generally, take a step towards developing transparent, personalized models that use speaker information in a controlled way.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- GlobalBench: A Benchmark for Global Progress in Natural Language ProcessingYueqi Song, Simran Khanuja, Pengfei Liu, Fahim Faisal 等EMNLP 2023 · 被引用 8 次
- KALM: Knowledge-Aware Integration of Local, Document, and Global Contexts for Long Document UnderstandingShangbin Feng, Zhaoxuan Tan, Wenqian Zhang, Zhenyu Lei 等ACL 2023 · 被引用 5 次
它引用的顶会 Paper11
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Deep Models Under the GAN: Information Leakage from Collaborative Deep LearningBriland Hitaj, Giuseppe Ateniese, Fernando Pérez-CruzCCS 2017 · 被引用 1,581 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 被引用 625 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
相关 Paper
- Discourse-Driven Code-Switching: Analyzing the Role of Content and Communicative Function in Spanish-English Bilingual SpeechDebasmita Bhattacharya, Juan Junco, Divya Tadimeti, Julia HirschbergEMNLP 2025
- Surprisal Predicts Code-Switching in Chinese-English Bilingual TextJesús Calvillo, Le Fang, Jeremy R. Cole, David ReitterEMNLP 2020 · 被引用 6 次
- From English to Code-Switching: Transfer Learning with Strong Morphological CluesGustavo Aguilar, Thamar SolorioACL 2020 · 被引用 1 次
- From Machine Translation to Code-Switching: Generating High-Quality Code-Switched TextIshan Tarunesh, Syamantak Kumar, Preethi JyothiACL 2021
- Code-switching Mediated Sentence-level Semantic LearningShuai Zhang, Jiangyan Yi, Zhengqi Wen, Jianhua Tao 等AAAI 2025
