DIWALI - Diversity and Inclusivity aWare cuLture specific Items for India: Dataset and Assessment of LLMs for Cultural Text Adaptation in Indian Context
Pramit Sahoo, Maharaj Brahma, Maunendra Sankar Desarkar
Abstract
Large language models (LLMs) are widely used in various tasks and applications.However, despite their wide capabilities, they are shown to lack cultural alignment (Ryan et al., 2024;AlKhamissi et al., 2024) and produce biased generations (Naous et al., 2024) due to a lack of cultural knowledge and competence.Evaluation of LLMs for cultural awareness and alignment is particularly challenging due to the lack of proper evaluation metrics and unavailability of culturally grounded datasets representing the vast complexity of cultures at the regional and sub-regional levels.Existing datasets for culture specific items (CSIs) focus primarily on concepts at the regional level and may contain false positives.To address this issue, we introduce a novel CSI dataset for Indian culture, belonging to 17 cultural facets.The dataset comprises 8k cultural concepts from 36 sub-regions.To measure the cultural competence of LLMs on a cultural text adaptation task, we evaluate the adaptations using the CSIs created, LLM as Judge, and human evaluations from diverse socio-demographic region.Furthermore, we perform quantitative analysis demonstrating selective sub-regional coverage and surface-level adaptations across all considered LLMs.Our dataset is available here: https://huggingface.co/datasets/nlip/DIWALI, project webpage 1 , and our codebase with model outputs can be found here: https://github.com/pramitsahoo/cultureevaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2cd7f55e-9c15-491d-8ea5-30f0c7025fdeCited by top-tier papers1
Ask how each one uses itBuilds on10
- Extracting Cultural Commonsense Knowledge at ScaleTuan-Phong Nguyen, Simon Razniewski, Aparna S. Varde, Gerhard WeikumWWW 2023 · 102 citations
- CultureLLM: Incorporating Cultural Differences into Large Language ModelsCheng Li, Mengzhuo Chen, Jindong Wang, Sunayana Sitaram et al.NeurIPS 2024 · 101 citations
- Language models are multilingual chain-of-thought reasonersFreda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang et al.ICLR 2023 · 52 citations
- Investigating Cultural Alignment of Large Language ModelsBadr AlKhamissi, Muhammad N. ElNokrashy, Mai Alkhamissi, Mona T. DiabACL 2024 · 27 citations
- Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language ModelsWenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai et al.ACL 2024 · 21 citations
Related papers
- SocialCC: Interactive Evaluation for Cultural Competence in Language AgentsJincenzi Wu, Jianxun Lian, Dingdong Wang, Helen M. MengACL 2025 · 8 citations
- Common to Whom? Regional Cultural Commonsense and LLM Bias in IndiaSangmitra Madhusudan, Trush Shashank More, Steph Buongiorno, Renata Dividino et al.ACL 2026 · 1 citation
- Culture-Aware Machine Translation in Large Language Models: Benchmarking and InvestigationZekun Yuan, Yangfan Ye, Xiaocheng Feng, Baohang Li et al.ACL 2026 · 2 citations
- DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian CultureArijit Maji, Raghvendra Kumar, Akash Ghosh, Anushka et al.EMNLP 2025 · 1 citation
- MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding EvaluationWeihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu et al.ACL 2026
