X-TURING: Towards an Enhanced and Efficient Turing Test for Long-Term Dialogue Agents
Weiqi Wu, Hongqiu Wu, Hai Zhao
Abstract
The Turing test examines whether AIs exhibit human-like behaviour in natural language conversations. The traditional setting limits each participant to one message at a time and requires constant human participation. This fails to reflect a natural conversational style and hinders the evaluation of dialogue agents based on Large Language Models (LLMs) in complex and prolonged interactions. This paper proposes X-Turing, which enhances the original test with a burst dialogue pattern, allowing more dynamic exchanges using consecutive messages. It further reduces human workload by iteratively generating dialogues that simulate the long-term interaction between the agent and a human to compose the majority of the test process. With the pseudo-dialogue history, the agent then engages in a shorter dialogue with a real human, which is paired with a human-human conversation on the same topic to be judged using questionnaires. We introduce the X-Turn Pass-Rate metric to assess the human likeness of LLMs across varying durations. While LLMs like GPT-4 initially perform well, achieving pass rates of 51.9% and 38.9% during 3 turns and 10 turns of dialogues respectively, their performance drops as the dialogue progresses, which underscores the difficulty in maintaining consistency in the long term.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a63a377c-2f8b-478b-a442-cdd073cfdefcCited by top-tier papers1
Ask how each one uses itBuilds on4
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Character-LLM: A Trainable Agent for Role-PlayingYunfan Shao, Linyang Li, Junqi Dai, Xipeng QiuEMNLP 2023 · 97 citations
- Large Language Models are Superpositions of All Characters: Attaining Arbitrary Role-play via Self-AlignmentKeming Lu, Bowen Yu, Chang Zhou, Jingren ZhouACL 2024 · 16 citations
- Large Scale Multi-Actor Generative Dialog ModelingAlex Boyd, Raul Puri, Mohammad Shoeybi, Mostofa Patwary et al.ACL 2020 · 1 citation
Related papers
- Human or Machine? A Preliminary Turing Test for Speech-to-Speech InteractionXiang Li, Jiabao Gao, Sipei Lin, Xuan Zhou et al.ICLR 2026 · 1 citation
- A Computational Framework for Evaluating Human-likeness in LLMs' Open-ended Human BehaviorsYuxuan Lei, Jianxun Lian, Defu Lian, Jincenzi Wu et al.ICML 2026
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject StudiesGati V. Aher, Rosa I. Arriaga, Adam Tauman KalaiICML 2023 · 651 citations
- IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question AnsweringRuosen Li, Ruochen Li, Barry Wang, Xinya DuNeurIPS 2024 · 26 citations
- Navigates Like Me: Understanding How People Evaluate Human-Like AI in Video GamesStephanie Milani, Arthur Juliani, Ida Momennejad, Raluca Georgescu et al.CHI 2023 · 17 citations
