Can LLMs Simulate L2-English Dialogue? An Information-Theoretic Analysis of L1-Dependent Biases
Rena Wei Gao, Xuetong Wu, Tatsuki Kuribayashi, Mingrui Ye, Siya Qi, Carsten Roever, Yuanxing Liu, Zheng Yuan, Jey Han Lau
Abstract
This study evaluates Large Language Models' (LLMs) ability to simulate non-native-like English use observed in human second language (L2) learners interfered with by their native first language (L1). In dialogue-based interviews, we prompt LLMs to mimic L2 English learners with specific L1s (e.g., Japanese, Thai, Urdu) across seven languages, comparing their outputs to real L2 learner data. Our analysis examines L1-driven linguistic biases, such as reference word usage and avoidance behaviors, using information-theoretic and distributional density measures. Results show that modern LLMs (e.g., Qwen2.5, LLAMA3, DeepseekV3, GPT 4o) replicate L1-dependent patterns observed in human L2 data, with distinct influences from various languages (e.g., Japanese, Korean, and Mandarin significantly affect tense agreement, and Urdu influences noun-verb collocations). Our results reveal LLMs' potential for L2 dialogue generation and evaluation for future educational applications. * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare QueriesYiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu et al.WWW 2024 · 126 citations
- Conversations Powered by Cross-Lingual KnowledgeWeiwei Sun, Chuan Meng, Qi Meng, Zhaochun Ren et al.SIGIR 2021 · 10 citations
- An Information-Theoretic Framework for Deep LearningHong Jun Jeon, Benjamin Van RoyNeurIPS 2022 · 8 citations
- SLABERT Talk Pretty One Day: Modeling Second Language Acquisition with BERTAditya Yadavalli, Alekhya Yadavalli, Vera TobinACL 2023 · 3 citations
Related papers
- Pride and Prejudice: LLM Amplifies Self-Bias in Self-RefinementWenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan et al.ACL 2024
- PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural DataIshaan Watts, Varun Gumma, Aditya Yadavalli, Vivek Seshadri et al.EMNLP 2024 · 3 citations
- Linguistic Bias in ChatGPT: Language Models Reinforce Dialect DiscriminationEve Fleisig, Genevieve Smith, Madeline Bossi, Ishita Rustagi et al.EMNLP 2024 · 36 citations
- Do Large Language Models have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMsYanzhu Guo, Simone Conia, Zelin Zhou, Min Li et al.ACL 2025
- Modeling Nonnative Sentence Processing with L2 Language ModelsTatsuya Aoyama, Nathan SchneiderEMNLP 2024 · 1 citation
