Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction
Xiang Li, Jiabao Gao, Sipei Lin, Xuan Zhou, Chi Zhang, Bo Cheng, Jiale Han, Benyou Wang
Abstract
The pursuit of human-like conversational agents has long been guided by the Turing test. For modern speech-to-speech (S2S) systems, a critical yet unanswered question is whether they can converse like humans. To tackle this, we conduct the first Turing test for S2S systems, collecting 2,968 human judgments on dialogues between 9 state-of-the-art S2S systems and 28 human participants. Our results deliver a clear finding: no existing evaluated S2S system passes the test, revealing a significant gap in human-likeness. To diagnose this failure, we develop a fine-grained taxonomy of 18 human-likeness dimensions and crowd-annotate our collected dialogues accordingly. Our analysis shows that the bottleneck is not semantic understanding but stems from paralinguistic features, emotional expressivity, and conversational persona. Furthermore, we find that off-the-shelf AI models perform unreliably as Turing test judges. In response, we propose an interpretable model that leverages the fine-grained human-likeness ratings and delivers accurate and transparent human-vs-machine discrimination, offering a powerful tool for automatic human-likeness evaluation. Our work 1 establishes the first humanlikeness evaluation for S2S systems and moves beyond binary outcomes to enable detailed diagnostic insights, paving the way for human-like improvements in conversational AI systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4bb02bd1-39b8-4371-a39c-6d730019e96cBuilds on3
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General IntelligenceSonal Kumar, Simon Sedlácek, Vaibhavi Lokegaonkar, Fernando López et al.AAAI 2026 · 1 citation
- Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face ConversationSe Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim et al.ACL 2024
Related papers
- Navigates Like Me: Understanding How People Evaluate Human-Like AI in Video GamesStephanie Milani, Arthur Juliani, Ida Momennejad, Raluca Georgescu et al.CHI 2023 · 17 citations
- S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech ModelsFeng Jiang, Zhiyu Lin, Yiyang Liu, Liumeng Xue et al.ACL 2026 · 17 citations
- ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech InteractionShu-Wen Yang, Ming Tu, Ting-Wei Liu, Xinghua Qu et al.ICLR 2026 · 29 citations
- X-TURING: Towards an Enhanced and Efficient Turing Test for Long-Term Dialogue AgentsWeiqi Wu, Hongqiu Wu, Hai ZhaoACL 2025 · 6 citations
- Towards Motion Turing Test: Evaluating Human-Likeness in Humanoid RobotsMingzhe Li, Mengyin Liu, Zekai Wu, Xincheng Lin et al.CVPR 2026 · 4 citations
