GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation
Mukul Khanna, Ram Ramrakhya, Gunjan Chhablani, Sriram Yenamandra, Théophile Gervet, Matthew Chang, Zsolt Kira, Devendra Singh Chaplot, Dhruv Batra, Roozbeh Mottaghi
摘要
mukulkhanna.github.io/goat-bench Figure 1 . We study the Go to Any Thing (GOAT) task, which involves agents navigating to a sequence of open vocabulary goals specified through any of the three modalities -category name, a language description, or an image. We propose GOAT-Bench, a benchmark for the GOAT task, where we evaluate modular and monolithic, explicit and implicit map-based navigation approaches. In the above example, we task the agent with sequentially navigating to 1) a recliner chair (from a closed set of k categories), 2) the oven shown in the picture, 3) "the white book on the coffee table in the living room", and some other objects in the scene. The goal of the benchmark is to facilitate progress towards building such universal, multi-modal, lifelong agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language NavigationXinda Xue, Junjun Hu, Minghua Luo, Xie Shichao 等ICLR 2026 · 被引用 51 次
- OctoNav: Towards Generalist Embodied NavigationChen Gao, Liankai Jin, Xingyu Peng, Jiazhao Zhang 等CVPR 2026 · 被引用 42 次
- 3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language ModelWenbo Hu, Yining Hong, Yanjun Wang, Leison Gao 等NeurIPS 2025 · 被引用 30 次
- AstraNav-Memory: Contexts Compression for Long MemoryJunjun Hu, Xinda Xue, Botao Ren, Minghua Luo 等CVPR 2026 · 被引用 5 次
- Multimodal LLM Guided Exploration and Active Mapping Using Fisher InformationWen Jiang, Boshu Lei, Katrina Ashton, Kostas DaniilidisICCV 2025 · 被引用 4 次
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang 等NeurIPS 2022 · 被引用 1,291 次
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 被引用 857 次
相关 Paper
- ZSON: Zero-Shot Object-Goal Navigation using Multimodal Goal EmbeddingsArjun Majumdar, Gunjan Aggarwal, Bhavika Devnani, Judy Hoffman 等NeurIPS 2022 · 被引用 344 次
- SOON: Scenario Oriented Object Navigation With Graph-Based ExplorationFengda Zhu, Xiwen Liang, Yi Zhu, Qizhi Yu 等CVPR 2021
- CoWs on Pasture: Baselines and Benchmarks for Language-Driven Zero-Shot Object NavigationSamir Yitzhak Gadre, Mitchell Wortsman, Gabriel Ilharco, Ludwig Schmidt 等CVPR 2023
- UINavBench: A Framework for Comprehensive Evaluation of Interactive Digital AgentsHarsh Agrawal, Eldon Schoop, Xinlei Pan, Anuj Mahajan 等ICCV 2025 · 被引用 9 次
- NavBench: Probing Multimodal Large Language Models for Embodied NavigationYanyuan Qiao, Haodong Hong, Wenqi Lyu, Dong An 等NeurIPS 2025 · 被引用 27 次
