WalkVLM: Aid Visually Impaired People Walking by Vision Language Model
Zhiqiang Yuan, Ting Zhang, Yeshuang Zhu, Jiapei Zhang, Ying Deng, Zexi Jia, Peixiang Luo, Xiaoyue Duan, Jie Zhou, Jinchao Zhang
摘要
Approximately 200 million individuals around the world suffer from varying degrees of visual impairment, making it crucial to leverage AI technology to offer walking assistance for these people. With the recent progress of visionlanguage models (VLMs), applying VLMs to offer walking guidance has become popular. However, the existing methods of walking guidance are mainly based on self-curated question-answering datasets that are not publicly accessible, without a standardized benchmark for training or evaluation. Moreover, walking assistance often requires realtime streaming video analysis and the generation of concise yet informative reminders, making VLMs struggle due to excessive responses and low efficiency in inferences. In this paper, we introduce the first large-scale dataset dedicated to walking assistance, comprising 12,000 video-annotation pairs, to provide a unified benchmark for training and evaluating systems to help visually-impaired individuals walk. Furthermore, a WalkVLM model is proposed, which employs chain of thought for hierarchical planning to generate concise but informative reminders and utilizes temporalaware adaptive prediction to reduce the temporal redundancy of reminders. Finally, we have established a solid benchmark for blind walking task and verified the advantages of WalkVLM in stream video processing for this task compared to other VLMs. Our dataset and code are available at the anonymous link https://walkvlm2024.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal LearningTianyi Bai, Yuxuan Fan, Jiantao Qiu, Fupeng Sun 等NeurIPS 2025 · 被引用 12 次
- GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware GuidanceJunhyeok Kim, Jaewoo Park, Junhee Park, Sangeyl Lee 等ACL 2026 · 被引用 7 次
- WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian NavigationRafi Ibn Sultan, Hui Zhu, Xiangyu Zhou, Chengyin Li 等CVPR 2026 · 被引用 4 次
- LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower RewardYi Zhao, Siqi Wang, Jing LiAAAI 2026 · 被引用 2 次
- Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired UsersAntonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas 等ACL 2025
它引用的顶会 Paper7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
相关 Paper
- Right this way: Can VLMs Guide Us to See More to Answer Questions?Li Liu, Diji Yang, Sijia Zhong, Kalyana Suma Sree Tholeti 等NeurIPS 2024 · 被引用 20 次
- VisAssist: A Visually Impaired-Captured Video Question Answering Benchmark for Assistive SystemsQi Gao, Heng Li, Yixin Zhou, Meixuan Zhou 等AAAI 2026
- VideoA11y: Method and Dataset for Accessible Video DescriptionChaoyu Li, Sid Padmanabhuni, Maryam S. Cheema, Hasti Seifi 等CHI 2025 · 被引用 23 次
- DR-VQA: Decompose-then-Reconstruct for Visual Question Answering in BLV AssistanceBocheng Pan, Hailong Shi, Xingyu GaoACM MM 2025
- IVQA-LD: Inclusive Multimodal Understanding for Population with Limb DeficiencyYan Ke, Xin Shen, Jiaying Ying, Xin Li 等ICML 2026
