One Step at a Time: Long-Horizon Vision-and-Language Navigation with Milestones
Chan Hee Song, Jihyung Kil, Tai-Yu Pan, Brian M. Sadler, Wei-Lun Chao, Yu Su
摘要
We study the problem of developing autonomous agents that can follow human instructions to infer and perform a sequence of actions to complete the underlying task. Significant progress has been made in recent years, especially for tasks with short horizons. However, when it comes to long-horizon tasks with extended sequences of actions, an agent can easily ignore some instructions or get stuck in the middle of the long instructions and eventually fail the task. To address this challenge, we propose a modelagnostic milestone-based task tracker (M-TRACK) to guide the agent and monitor its progress. Specifically, we propose a milestone builder that tags the instructions with navigation and interaction milestones which the agent needs to complete step by step, and a milestone checker that systemically checks the agent's progress in its current milestone and determines when to proceed to the next. On the challenging ALFRED dataset, our M-TRACK leads to a notable 33% and 52% relative improvement in unseen success rate over two competitive base models. Check our code at https://github.com/chanhee-luke/M-Track .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao 等ICCV 2023 · 被引用 685 次
- Do Embodied Agents Dream of Pixelated Sheep: Embodied Decision Making using Language Guided World ModellingKolby Nottingham, Prithviraj Ammanabrolu, Alane Suhr, Yejin Choi 等ICML 2023 · 被引用 110 次
- Learning Vision-and-Language Navigation from YouTube VideosKunyang Lin, Peihao Chen, Diwei Huang, Thomas H. Li 等ICCV 2023 · 被引用 57 次
- Don't Generate, Discriminate: A Proposal for Grounding Language Models to Real-World EnvironmentsYu Gu, Xiang Deng, Yu SuACL 2023 · 被引用 36 次
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsChun-Hsiao Yeh, Chenyu Wang, Shengbang Tong, Ta Ying Cheng 等AAAI 2026 · 被引用 35 次
它引用的顶会 Paper13
- Episodic Transformer for Vision-and-Language NavigationAlexander Pashevich, Cordelia Schmid, Chen SunICCV 2021 · 被引用 228 次
- FILM: Following Instructions in Language with Modular MethodsSo Yeon Min, Devendra Singh Chaplot, Pradeep Kumar Ravikumar, Yonatan Bisk 等ICLR 2022 · 被引用 189 次
- Language and Visual Entity Relationship Graph for Agent NavigationYicong Hong, Cristian Rodriguez Opazo, Yuankai Qi, Qi Wu 等NeurIPS 2020 · 被引用 167 次
- Sub-Instruction Aware Vision-and-Language NavigationYicong Hong, Cristian Rodriguez Opazo, Qi Wu, Stephen GouldEMNLP 2020 · 被引用 55 次
- Factorizing Perception and Policy for Interactive Instruction FollowingKunal Pratap Singh, Suvaansh Bhambri, Byeonghwi Kim, Roozbeh Mottaghi 等ICCV 2021 · 被引用 39 次
相关 Paper
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday TasksMohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk 等CVPR 2020
- Milestone-Guided Policy Learning for Long-Horizon Language AgentsZixuan Wang, Yuchen Yan, Hongxing Li, Teng Pan 等ICML 2026 · 被引用 8 次
- BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby StepsWang Zhu, Hexiang Hu, Jiacheng Chen, Zhiwei Deng 等ACL 2020 · 被引用 62 次
- SkillTracer: Structural Failure Attribution and Refinement of Agentic Skills in Long-Horizon Web TasksYuyang Li, Yiran Dou, Jie-Jing Shao, Yueming Lyu 等KDD 2026 · 被引用 5 次
- RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World UsersSuyu Ye, Haojun Shi, Darren Shih, Hyokun Yun 等AAAI 2026 · 被引用 17 次
