Navigation Turing Test (NTT): Learning to Evaluate Human-Like Navigation
Sam Devlin, Raluca Georgescu, Ida Momennejad, Jaroslaw Rzepecki, Evelyn Zuniga, Gavin Costello, Guy Leroy, Ali Shaw, Katja Hofmann
摘要
A key challenge on the path to developing agents that learn complex human-like behavior is the need to quickly and accurately quantify human-likeness. While human assessments of such behavior can be highly accurate, speed and scalability are limited. We address these limitations through a novel automated Navigation Turing Test (ANTT) that learns to predict human judgments of human-likeness. We demonstrate the effectiveness of our automated NTT on a navigation task in a complex 3D environment. We investigate six classification models to shed light on the types of architectures best suited to this task, and validate them against data collected through a human NTT. Our best models achieve high accuracy when distinguishing true human and agent behavior. At the same time, we show that predicting finer-grained human assessment of agents' progress towards human-like behavior remains unsolved. Our work takes an important step towards agents that more effectively learn complex human-like behavior.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Navigates Like Me: Understanding How People Evaluate Human-Like AI in Video GamesStephanie Milani, Arthur Juliani, Ida Momennejad, Raluca Georgescu 等CHI 2023 · 被引用 17 次
- Beyond Accuracy: Tracking more like Human via Visual SearchDailing Zhang, Shiyu Hu, Xiaokun Feng, Xuchen Li 等NeurIPS 2024 · 被引用 8 次
- Learning Human-Like RL Agents Through Trajectory Optimization With Action QuantizationJian-Ting Guo, Yu-Cheng Chen, Ping-Chun Hsieh, Kuo-Hao Ho 等NeurIPS 2025 · 被引用 3 次
- Playing the Imitation Game: How Perceived Generated Content Shapes Player ExperienceMahsa Bazzaz, Seth CooperCHI 2026
它引用的顶会 Paper3
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 被引用 857 次
- Learning to Contextually Aggregate Multi-Source Supervision for Sequence LabelingOuyu Lan, Xiao Huang, Bill Yuchen Lin, He Jiang 等ACL 2020 · 被引用 33 次
- Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied NavigationJuncheng Li, Xin Wang, Siliang Tang, Haizhou Shi 等CVPR 2020
相关 Paper
- Towards Motion Turing Test: Evaluating Human-Likeness in Humanoid RobotsMingzhe Li, Mengyin Liu, Zekai Wu, Xincheng Lin 等CVPR 2026 · 被引用 4 次
- Human or Machine? A Preliminary Turing Test for Speech-to-Speech InteractionXiang Li, Jiabao Gao, Sipei Lin, Xuan Zhou 等ICLR 2026 · 被引用 1 次
- A Computational Framework for Evaluating Human-likeness in LLMs' Open-ended Human BehaviorsYuxuan Lei, Jianxun Lian, Defu Lian, Jincenzi Wu 等ICML 2026
- A Newborn Embodied Turing Test for Comparing Object Segmentation Across Animals and MachinesManju Garimella, Denizhan Pak, Justin N. Wood, Samantha Marie Waters WoodICLR 2024 · 被引用 1 次
- SPACeR: Self-Play Anchoring with Centralized Reference ModelsWei-Jer Chang, Akshay Rangesh, Kevin Joseph, Matthew Strong 等ICLR 2026 · 被引用 9 次
