Navigation Turing Test (NTT): Learning to Evaluate Human-Like Navigation
Sam Devlin, Raluca Georgescu, Ida Momennejad, Jaroslaw Rzepecki, Evelyn Zuniga, Gavin Costello, Guy Leroy, Ali Shaw, Katja Hofmann
Abstract
A key challenge on the path to developing agents that learn complex human-like behavior is the need to quickly and accurately quantify human-likeness. While human assessments of such behavior can be highly accurate, speed and scalability are limited. We address these limitations through a novel automated Navigation Turing Test (ANTT) that learns to predict human judgments of human-likeness. We demonstrate the effectiveness of our automated NTT on a navigation task in a complex 3D environment. We investigate six classification models to shed light on the types of architectures best suited to this task, and validate them against data collected through a human NTT. Our best models achieve high accuracy when distinguishing true human and agent behavior. At the same time, we show that predicting finer-grained human assessment of agents' progress towards human-like behavior remains unsolved. Our work takes an important step towards agents that more effectively learn complex human-like behavior.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f990a23-12ef-4c29-a16d-99e2328e43bbCited by top-tier papers4
- Navigates Like Me: Understanding How People Evaluate Human-Like AI in Video GamesStephanie Milani, Arthur Juliani, Ida Momennejad, Raluca Georgescu et al.CHI 2023 · 17 citations
- Beyond Accuracy: Tracking more like Human via Visual SearchDailing Zhang, Shiyu Hu, Xiaokun Feng, Xuchen Li et al.NeurIPS 2024 · 8 citations
- Learning Human-Like RL Agents Through Trajectory Optimization With Action QuantizationJian-Ting Guo, Yu-Cheng Chen, Ping-Chun Hsieh, Kuo-Hao Ho et al.NeurIPS 2025 · 3 citations
- Playing the Imitation Game: How Perceived Generated Content Shapes Player ExperienceMahsa Bazzaz, Seth CooperCHI 2026
Builds on3
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 857 citations
- Learning to Contextually Aggregate Multi-Source Supervision for Sequence LabelingOuyu Lan, Xiao Huang, Bill Yuchen Lin, He Jiang et al.ACL 2020 · 33 citations
- Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied NavigationJuncheng Li, Xin Wang, Siliang Tang, Haizhou Shi et al.CVPR 2020
Related papers
- Towards Motion Turing Test: Evaluating Human-Likeness in Humanoid RobotsMingzhe Li, Mengyin Liu, Zekai Wu, Xincheng Lin et al.CVPR 2026 · 4 citations
- Human or Machine? A Preliminary Turing Test for Speech-to-Speech InteractionXiang Li, Jiabao Gao, Sipei Lin, Xuan Zhou et al.ICLR 2026 · 1 citation
- A Computational Framework for Evaluating Human-likeness in LLMs' Open-ended Human BehaviorsYuxuan Lei, Jianxun Lian, Defu Lian, Jincenzi Wu et al.ICML 2026
- A Newborn Embodied Turing Test for Comparing Object Segmentation Across Animals and MachinesManju Garimella, Denizhan Pak, Justin N. Wood, Samantha Marie Waters WoodICLR 2024 · 1 citation
- SPACeR: Self-Play Anchoring with Centralized Reference ModelsWei-Jer Chang, Akshay Rangesh, Kevin Joseph, Matthew Strong et al.ICLR 2026 · 9 citations
