Large Language Models Discriminate Against Speakers of German Dialects
Minh Duc Bui, Carolin Holtermann, Valentin Hofmann, Anne Lauscher, Katharina von der Wense
Abstract
Dialects represent a significant component of human culture and are found across all regions of the world. In Germany, more than 40% of the population speaks a regional dialect (Adler and Hansen, 2022) . However, despite cultural importance, individuals speaking dialects often face negative societal stereotypes. We examine whether such stereotypes are mirrored by large language models (LLMs). We draw on the sociolinguistic literature on dialect perception to analyze traits commonly associated with dialect speakers. Based on these traits, we assess the dialect naming bias and dialect usage bias expressed by LLMs in two tasks: an association task and a decision task. To assess a model's dialect usage bias, we construct a novel evaluation corpus that pairs sentences from seven regional German dialects (e.g., Alemannic and Bavarian) with their standard German counterparts. We find that: (1) in the association task, all evaluated LLMs exhibit significant dialect naming and dialect usage bias against German dialect speakers, reflected in negative adjective associations; (2) all models reproduce these dialect naming and dialect usage biases in their decision making; and (3) contrary to prior work showing minimal bias with explicit demographic mentions, we find that explicitly labeling linguistic demographics-German dialect speakers-amplifies bias more than implicit cues like dialect usage. * Equal contribution. 1 The literature disagrees on an exact definition; we give more information in Appendix A.1. We examine whether LLMs exhibit the same dialect-related stereotypes found in humans. Building on prior work in dialect perception (Gärtig et al., 2010; Trillhaase, 2021), we concentrate on stereo-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc4a532c-0383-46ed-b839-8977e6da8272Builds on3
- Crosslingual Generalization through Multitask FinetuningNiklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts et al.ACL 2023 · 319 citations
- Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language ModelsMyra Cheng, Esin Durmus, Dan JurafskyACL 2023 · 89 citations
- Linguistic Bias in ChatGPT: Language Models Reinforce Dialect DiscriminationEve Fleisig, Genevieve Smith, Madeline Bossi, Ishita Rustagi et al.EMNLP 2024 · 36 citations
Related papers
- Are Stereotypes Leading LLMs' Zero-Shot Stance Detection ?Anthony Dubreuil, Antoine Gourru, Christine Largeron, Amine TrabelsiEMNLP 2025 · 1 citation
- Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM CollaborationWeicheng Ma, John J. Guerrerio, Soroush VosoughiEMNLP 2025 · 1 citation
- Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language ModelsZara Siddique, Liam D. Turner, Luis Espinosa AnkeEMNLP 2024 · 2 citations
- AccessEval: Benchmarking Disability Bias in Large Language ModelsSrikant Panda, Amit Agarwal, Hitesh Laxmichand PatelEMNLP 2025 · 2 citations
- Deus Ex Machina and Personas from Large Language Models: Investigating the Composition of AI-Generated Persona DescriptionsJoni Salminen, Chang Liu, Wenjing Pian, Jianxing Chi et al.CHI 2024 · 55 citations
