Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective Responses
Chongyuan Dai, Yaling Shen, Zihan Gao, Jia Li, Yishun Jiang, Yaxiong Wang, Liu Liu, Zongyuan Ge, Jinpeng Hu
Abstract
Culture serves as a fundamental determinant of human affective processing and profoundly shapes how individuals perceive and interpret emotional stimuli. Despite this intrinsic link extant evaluations regarding cultural alignment within Large Language Models primarily prioritize declarative knowledge such as geographical facts or established societal customs. These benchmarks remain insufficient to capture the subjective interpretative variance inherent to diverse sociocultural lenses. To address this limitation, we introduce CEDAR, a multimodal benchmark constructed entirely from scenarios capturing Culturally Elicited Distinct Affective Responses. To construct CEDAR, we implement a novel pipeline that leverages LLMgenerated provisional labels to isolate instances yielding cross-cultural emotional distinctions, and subsequently derives reliable ground-truth annotations through rigorous human evaluation. The resulting benchmark comprises 10,962 instances across seven languages and 14 finegrained emotion categories, with each language including 400 multimodal and 1,166 text-only samples. Comprehensive evaluations of 17 representative multilingual models reveal a dissociation between language consistency and cultural alignment, demonstrating that culturally grounded affective understanding remains a significant challenge for current models. Codes and datasets will be released soon 1 . * Equal contribution 1 https://github.com/MindIntLab-HFUT/CEDAR ar How does the person standing near the lanterns feel during this moment? MULTIMODAL Guests fill the room, knowing the event celebrates new beginnings. Before arrival, one guest wraps a gift in black paper. You hesitate, unsure how to react. How do you feel upon seeing the black-wrapped gift?
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29712f01-b7ae-47da-bb94-521144687cbaBuilds on22
- Knowledge of cultural moral norms in large language modelsAida Ramezani, Yang XuACL 2023 · 44 citations
- Benchmarking Vision Language Models for Cultural UnderstandingShravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy et al.EMNLP 2024 · 26 citations
- Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language ModelsWenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai et al.ACL 2024 · 21 citations
- Commonsense Reasoning in Arab CultureAbdelrahman Boda Sadallah, Junior Cedric Tonga, Khalid Almubarak, Saeed Almheiri et al.ACL 2025 · 18 citations
- ArtELingo: A Million Emotion Annotations of WikiArt with Emphasis on Diversity over Language and CultureYoussef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Feifan Li et al.EMNLP 2022 · 13 citations
Related papers
- MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding EvaluationWeihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu et al.ACL 2026
- HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized BenchmarksTing Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding et al.CVPR 2026
- AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language ModelsZheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun et al.ICML 2025
- CULEMO: Cultural Lenses on Emotion - Benchmarking LLMs for Cross-Cultural Emotion UnderstandingTadesse Destaw Belay, Ahmed Haj Ahmed, Alvin Grissom II, Iqra Ameer et al.ACL 2025
- EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional UnderstandingPengze Guo, Jingxi Liang, Zhiwen Xie, Qifeng Wang et al.ACL 2026
