Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
Eunsu Kim, Junyeong Park, Juhyun Oh, Kiwoong Park, Seyoung Song, A. Seza Dogruöz, Alice Oh, Najoung Kim
Abstract
As LLMs are increasingly deployed in real-world interactions, their social reasoning in interpersonal communication becomes critical. To explore their capabilities, we introduce SCRIPTS, a 1.1k-dialogue dataset in English and Korean, sourced from movie scripts and propose a social reasoning task based on SCRIPTS that evaluates the capacity of LLMs to infer the social relationships (e.g., friends, lovers) between speakers in each dialogue. Evaluating nine models on our task, current LLMs achieve around 75--80% on the English dataset and 58--69% in Korean, and models predict an Unlikely relationship in 10--25% of responses in both languages. Furthermore, we find that thinking models and chain-of-thought prompting provide minimal benefits for social reasoning and occasionally amplify social biases. In sum, there are significant limitations in current LLMs'social reasoning capabilities, especially for Korean, highlighting the need for efforts to develop socially-aware LLMs across languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 258bda48-147a-457a-8667-15db1acfff15Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- DDRel: A New Dataset for Interpersonal Relation Classification in Dyadic DialoguesQi Jia, Hongru Huang, Kenny Q. ZhuAAAI 2021 · 23 citations
- Your spouse needs professional help: Determining the Contextual Appropriateness of Messages through Modeling Social RelationshipsDavid Jurgens, Agrima Seth, Jackson Sargent, Athena Aghighi et al.ACL 2023 · 4 citations
- PRIDE: Predicting Relationships in ConversationsAnna Tigunova, Paramita Mirza, Andrew Yates, Gerhard WeikumEMNLP 2021
Related papers
- EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product AssociationWeiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag et al.ACL 2025 · 15 citations
- Read the Room: Video Social Reasoning with Mental-Physical Causal ChainsLixing Niu, Jiapeng Li, Xingping Yu, Xinyi Dong et al.ICLR 2026
- Reasoning over Uncertain Text by Generative Large Language ModelsAliakbar Nafar, Kristen Brent Venable, Parisa KordjamshidiAAAI 2025 · 13 citations
- RecToM: A Benchmark for Evaluating Machine Theory of Mind in LLM-based Conversational Recommender SystemsMengfan Li, Xuanhua Shi, Yang DengAAAI 2026
- Eliciting Better Multilingual Structured Reasoning from LLMs through CodeBryan Li, Tamer Alkhouli, Daniele Bonadiman, Nikolaos Pappas et al.ACL 2024
