CAReDiO: Enhancing Cultural Alignment of LLM via Representativeness and Distinctiveness Guided Data Optimization
Jing Yao, Xiaoyuan Yi, Jindong Wang, Zhicheng Dou, Xing Xie
Abstract
As Large Language Models (LLMs) more deeply integrate into human life across various regions, aligning them with pluralistic cultures is crucial for improving user engagement and mitigating cultural conflicts. For this purpose, recently, different culture-specific corpora have been carefully curated, either synthesized or manually annotated. Nevertheless, inspired by culture theories, we identify two key challenges faced by these datasets: (1) Representativeness: These corpora fail to fully capture the target culture's core characteristics, causing insufficient cultural coverage with redundancy; (2) Distinctiveness: They struggle to distinguish the unique nuances of a given culture from shared patterns across other relevant ones, hindering precise cultural modelling. To handle these challenges, we introduce CAReDiO, a novel data optimization framework, which alternatively refines culture-sensitive questions and responses according to information-theoretic objectives in an in-context optimization manner, enhancing the cultural informativeness and distinguishability of constructed data. Extensive experiments on 15 distinct cultures demonstrate that CAReDiO can create high-quality data with richer cultural information and enable efficient alignment of small open-source or large proprietary LLMs with as few as 200 training samples, consistently outperforming previous datasets in both multi-choice and open-ended cultural benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e6d1249-36b6-42e3-8177-220a5fbe5facCited by top-tier papers2
- Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language ModelsHanze Guo, Jing Yao, Xiao Zhou, Xiaoyuan Yi et al.NeurIPS 2025 · 6 citations
- An Experimental Study on the Influence of Culture on Cross-Lingual Sentiment TransferAhao Liu, Haitong Yang, Chuanrong WangACL 2026
Builds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- Generalized Jensen-Shannon Divergence Loss for Learning with Noisy LabelsErik Englesson, Hossein AzizpourNeurIPS 2021 · 170 citations
- Extracting Cultural Commonsense Knowledge at ScaleTuan-Phong Nguyen, Simon Razniewski, Aparna S. Varde, Gerhard WeikumWWW 2023 · 102 citations
Related papers
- CARE: Multilingual Human Preference Learning for Cultural AwarenessGeyang Guo, Tarek Naous, Hiromi Wakaki, Yukiko Nishimura et al.EMNLP 2025
- Mind the Gap in Cultural Alignment: Task-Aware Culture Management for Large Language ModelsBinchi Zhang, Xujiang Zhao, Jundong Li, Haifeng Chen et al.ACL 2026 · 3 citations
- CulturePark: Boosting Cross-cultural Understanding in Large Language ModelsCheng Li, Damien Teney, Linyi Yang, Qingsong Wen et al.NeurIPS 2024 · 34 citations
- Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic PromptingSagnik Mukherjee, Muhammad Farid Adilazuarda, Sunayana Sitaram, Kalika Bali et al.EMNLP 2024 · 3 citations
- CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data SynthesisRuixiang Feng, Shen Gao, Xiuying Chen, Lisi Chen et al.ACL 2025
