Women, Infamous, and Exotic Beings: A Comparative Study of Honorific Usages in Wikipedia and LLMs for Bengali and Hindi
Sourabrata Mukherjee, Atharva Mehta, Sougata Saha, Akhil Arora, Monojit Choudhury
摘要
The obligatory use of third-person honorifics is a distinctive feature of several South Asian languages, encoding nuanced socio-pragmatic cues such as power, age, gender, fame, and social distance. In this work, (i) We present the first large-scale study of third-person honorific pronoun and verb usage across 10,000 Hindi and Bengali Wikipedia articles with annotations linked to key socio-demographic attributes of the subjects, including gender, age group, fame, and cultural origin. (ii) Our analysis uncovers systematic intra-language regularities but notable cross-linguistic differences: honorifics are more prevalent in Bengali than in Hindi, while non-honorifics dominate while referring to infamous, juvenile, and culturally "exotic" entities. Notably, in both languages, and more prominently in Hindi, men are more frequently addressed with honorifics than women. (iii) To examine whether large language models (LLMs) internalize similar socio-pragmatic norms, we probe six LLMs using controlled generation and translation tasks over 1,000 culturally balanced entities. We find that LLMs diverge from Wikipedia usage, exhibiting alternative preferences in honorific selection across tasks, languages, and sociodemographic attributes. These discrepancies highlight gaps in the socio-cultural alignment of LLMs and open new directions for studying how LLMs acquire, adapt, or distort sociallinguistic norms. Our code and data are publicly available at https://github.com/souro/ honorific-wiki-llm
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological GenerationMehul Agarwal, Aditya Aggarwal, Arnav Goel, Medha Hira 等ACL 2026
- Do LLM Agents Mirror Socio-Cognitive Effects in Power-Asymmetric Conversations?Anvesh Rao Vijjini, Sagar Manjunath, Snigdha ChaturvediACL 2026
- An Empirical Analysis of the Writing Styles of Persona-Assigned LLMsManuj Malik, Jing Jiang, Kian Ming A. ChaiEMNLP 2024 · 被引用 2 次
- Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion AttributionFlor Miriam Plaza del Arco, Amanda Cercas Curry, Alba Cercas Curry, Gavin Abercrombie 等ACL 2024 · 被引用 8 次
- The Cross-linguistic Role of Animacy in Grammar StructuresNina Gregorio, Matteo Gay, Sharon Goldwater, Edoardo M. PontiACL 2025
