DIWALI - Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Context
Pramit Sahoo, Maharaj Brahma, Maunendra Sankar Desarkar
摘要
Large language models (LLMs) are widely used in various tasks and applications.However, despite their wide capabilities, they are shown to lack cultural alignment (Ryan et al., 2024;AlKhamissi et al., 2024) and produce biased generations (Naous et al., 2024) due to a lack of cultural knowledge and competence.Evaluation of LLMs for cultural awareness and alignment is particularly challenging due to the lack of proper evaluation metrics and unavailability of culturally grounded datasets representing the vast complexity of cultures at the regional and sub-regional levels.Existing datasets for culture specific items (CSIs) focus primarily on concepts at the regional level and may contain false positives.To address this issue, we introduce a novel CSI dataset for Indian culture, belonging to 17 cultural facets.The dataset comprises 8k cultural concepts from 36 sub-regions.To measure the cultural competence of LLMs on a cultural text adaptation task, we evaluate the adaptations using the CSIs created, LLM as Judge, and human evaluations from diverse socio-demographic region.Furthermore, we perform quantitative analysis demonstrating selective sub-regional coverage and surface-level adaptations across all considered LLMs.Our dataset is available here: https://huggingface.co/datasets/nlip/DIWALI, project webpage 1 , and our codebase with model outputs can be found here: https://github.com/pramitsahoo/cultureevaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Extracting Cultural Commonsense Knowledge at ScaleTuan-Phong Nguyen, Simon Razniewski, Aparna S. Varde, Gerhard WeikumWWW 2023 · 被引用 102 次
- CultureLLM: Incorporating Cultural Differences into Large Language ModelsCheng Li, Mengzhuo Chen, Jindong Wang, Sunayana Sitaram 等NeurIPS 2024 · 被引用 101 次
- Language models are multilingual chain-of-thought reasonersFreda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang 等ICLR 2023 · 被引用 52 次
- Investigating Cultural Alignment of Large Language ModelsBadr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. DiabACL 2024 · 被引用 27 次
- Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language ModelsWenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai 等ACL 2024 · 被引用 21 次
相关 Paper
- SocialCC: Interactive Evaluation for Cultural Competence in Language AgentsJincenzi Wu, Jianxun Lian, Dingdong Wang, Helen M. MengACL 2025 · 被引用 8 次
- Common to Whom? Regional Cultural Commonsense and LLM Bias in IndiaSangmitra Madhusudan, Trush Shashank More, Steph Buongiorno, Renata Dividino 等ACL 2026 · 被引用 1 次
- Culture-Aware Machine Translation in Large Language Models: Benchmarking and InvestigationZekun Yuan, Yangfan Ye, Xiaocheng Feng, Baohang Li 等ACL 2026 · 被引用 2 次
- DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian CultureArijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka 等EMNLP 2025 · 被引用 1 次
- MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding EvaluationWeihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu 等ACL 2026
