Can LLMs Simulate L2-English Dialogue? An Information-Theoretic Analysis of L1-Dependent Biases
Rena Wei Gao, Xuetong Wu, Tatsuki Kuribayashi, Mingrui Ye, Siya Qi, Carsten Roever, Yuanxing Liu, Zheng Yuan, Jey Han Lau
摘要
This study evaluates Large Language Models' (LLMs) ability to simulate non-native-like English use observed in human second language (L2) learners interfered with by their native first language (L1). In dialogue-based interviews, we prompt LLMs to mimic L2 English learners with specific L1s (e.g., Japanese, Thai, Urdu) across seven languages, comparing their outputs to real L2 learner data. Our analysis examines L1-driven linguistic biases, such as reference word usage and avoidance behaviors, using information-theoretic and distributional density measures. Results show that modern LLMs (e.g., Qwen2.5, LLAMA3, DeepseekV3, GPT 4o) replicate L1-dependent patterns observed in human L2 data, with distinct influences from various languages (e.g., Japanese, Korean, and Mandarin significantly affect tense agreement, and Urdu influences noun-verb collocations). Our results reveal LLMs' potential for L2 dialogue generation and evaluation for future educational applications. * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare QueriesYiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu 等WWW 2024 · 被引用 126 次
- Conversations Powered by Cross-Lingual KnowledgeWeiwei Sun, Chuan Meng, Qi Meng, Zhaochun Ren 等SIGIR 2021 · 被引用 10 次
- An Information-Theoretic Framework for Deep LearningHong Jun Jeon, Benjamin Van RoyNeurIPS 2022 · 被引用 8 次
- SLABERT Talk Pretty One Day: Modeling Second Language Acquisition with BERTAditya Yadavalli, Alekhya Yadavalli, Vera TobinACL 2023 · 被引用 3 次
相关 Paper
- Pride and Prejudice: LLM Amplifies Self-Bias in Self-RefinementWenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan 等ACL 2024
- PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural DataIshaan Watts, Varun Gumma, Aditya Yadavalli, Vivek Seshadri 等EMNLP 2024 · 被引用 3 次
- Linguistic Bias in ChatGPT: Language Models Reinforce Dialect DiscriminationEve Fleisig, Genevieve Smith, Madeline Bossi, Ishita Rustagi 等EMNLP 2024 · 被引用 36 次
- Do Large Language Models have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMsYanzhu Guo, Simone Conia, Zelin Zhou, Min Li 等ACL 2025
- Modeling Nonnative Sentence Processing with L2 Language ModelsTatsuya Aoyama, Nathan SchneiderEMNLP 2024 · 被引用 1 次
