Evaluating Theory of (an uncertain) Mind: Predicting the Uncertain Beliefs of Others from Conversational Cues
Anthony B. Sicilia, Malihe Alikhani
摘要
Typically, when evaluating Theory of Mind, we consider the beliefs of others to be binary: held or not held. But what if someone is unsure about their own beliefs? How can we quantify this uncertainty? We propose a new suite of tasks, challenging language models (LMs) to model the uncertainty of participants in a dialogue. We design these tasks around conversation forecasting, where the goal is to predict the probability of an unobserved conversation outcome. Uniquely, we view conversation agents themselves as forecasters, asking an LM to predict the uncertainty of an individual from their language use. We experiment with scaling methods, bagging, and demographic context for this regression task, conducting experiments on three dialogue corpora (social, negotiation, task-oriented) with eight LMs. While LMs can explain up to 7% variance in the uncertainty of others, we highlight the difficulty of the tasks and room for future work, especially in tasks that require explicit shifts in perspective. How certain is S1 they are more satisfied than would occur by chance? → Ground-truth P = 11% → Predicted P = 40% How certain is S2 they like S1 more than would occur by chance? → Ground-truth P = 27% → Predicted P = 90% How certain is S2 they are more satisfied than would occur by chance? → Ground-truth P = 81% → Predicted P = 90%
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMsMaarten Sap, Ronan Le Bras, Daniel Fried, Yejin ChoiEMNLP 2022 · 被引用 92 次
- Monarch Mixer: A Simple Sub-Quadratic GEMM-Based ArchitectureDaniel Y. Fu, Simran Arora, Jessica Grogan, Isys Johnson 等NeurIPS 2023 · 被引用 80 次
- Theory of Mind for Multi-Agent Collaboration via Large Language ModelsHuao Li, Yu Quan Chong, Simon Stepputtis, Joseph Campbell 等EMNLP 2023 · 被引用 57 次
相关 Paper
- Perceptions of Linguistic Uncertainty by Language Models and HumansCatarina G. Belém, Markelle Kelly, Mark Steyvers, Sameer Singh 等EMNLP 2024 · 被引用 6 次
- SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?Michael Kirchhof, Luca Füger, Adam Golinski, Eeshan Gunesh Dhekane 等ICLR 2026 · 被引用 4 次
- OpenEstimate: Evaluating LLMs on Reasoning Under Uncertainty with Real-World DataAlana Renda, Jillian Ross, Jacob AndreasICLR 2026 · 被引用 3 次
- AutoToM: Scaling Model-based Mental Inference via Automated Agent ModelingZhining Zhang, Chuanyang Jin, Mung Yao Jia, Shunchi Zhang 等NeurIPS 2025 · 被引用 30 次
- RecToM: A Benchmark for Evaluating Machine Theory of Mind in LLM-based Conversational Recommender SystemsMengfan Li, Xuanhua Shi, Yang DengAAAI 2026
