GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation
Mukul Khanna, Ram Ramrakhya, Gunjan Chhablani, Sriram Yenamandra, Théophile Gervet, Matthew Chang, Zsolt Kira, Devendra Singh Chaplot, Dhruv Batra, Roozbeh Mottaghi
Abstract
mukulkhanna.github.io/goat-bench Figure 1 . We study the Go to Any Thing (GOAT) task, which involves agents navigating to a sequence of open vocabulary goals specified through any of the three modalities -category name, a language description, or an image. We propose GOAT-Bench, a benchmark for the GOAT task, where we evaluate modular and monolithic, explicit and implicit map-based navigation approaches. In the above example, we task the agent with sequentially navigating to 1) a recliner chair (from a closed set of k categories), 2) the oven shown in the picture, 3) "the white book on the coffee table in the living room", and some other objects in the scene. The goal of the benchmark is to facilitate progress towards building such universal, multi-modal, lifelong agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language NavigationXinda Xue, Junjun Hu, Minghua Luo, Xie Shichao et al.ICLR 2026 · 51 citations
- OctoNav: Towards Generalist Embodied NavigationChen Gao, Liankai Jin, Xingyu Peng, Jiazhao Zhang et al.CVPR 2026 · 42 citations
- 3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language ModelWenbo Hu, Yining Hong, Yanjun Wang, Leison Gao et al.NeurIPS 2025 · 30 citations
- AstraNav-Memory: Contexts Compression for Long MemoryJunjun Hu, Xinda Xue, Botao Ren, Minghua Luo et al.CVPR 2026 · 5 citations
- Multimodal LLM Guided Exploration and Active Mapping Using Fisher InformationWen Jiang, Boshu Lei, Katrina Ashton, Kostas DaniilidisICCV 2025 · 4 citations
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 857 citations
Related papers
- ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal EmbeddingsArjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman et al.NeurIPS 2022 · 344 citations
- SOON: Scenario Oriented Object Navigation With Graph-Based ExplorationFengda Zhu, Xiwen Liang, Yi Zhu, Qizhi Yu et al.CVPR 2021
- CoWs on Pasture: Baselines and Benchmarks for Language-Driven Zero-Shot Object NavigationSamir Yitzhak Gadre, Mitchell Wortsman, Gabriel Ilharco, Ludwig Schmidt et al.CVPR 2023
- UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital AgentsHarsh Agrawal, Eldon Schoop, Xinlei Pan, Anuj Mahajan et al.ICCV 2025 · 9 citations
- NavBench: Probing Multimodal Large Language Models for Embodied NavigationYanyuan Qiao, Haodong Hong, Wenqi Lyu, Dong An et al.NeurIPS 2025 · 27 citations
