GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
Junhyeok Kim, Jaewoo Park, Junhee Park, Sangeyl Lee, Jiwan Chung, Jisung Kim, Ji Hoon Joung, Youngjae Yu
Abstract
For people affected by blindness and low vision (BLV), safe and independent navigation remains a major challenge, impacting over 2.2 billion individuals worldwide. Although multimodal large language models (MLLMs) offer new opportunities for assistive navigation, progress has been limited by the scarcity of accessibility-aware datasets, because creating them requires labor-intensive expert annotation. To this end, we introduce GUIDEDOG, a novel dataset containing 22K image-description pairs (2K human-verified) capturing real-world pedestrian scenes across 46 countries. Our human-AI pipeline shifts annotation from generation to verification, grounded in established BLV guidance standards from experts and research, improving scalability while maintaining quality. We also present GUIDEDOGQA, an 818-sample benchmark evaluating object recognition and depth perception. Experiments reveal that depth perception and adherence to these standards remain challenging for current MLLMs. * Equal Contribution. † Work done while at SK Telecom. S2 S2 S1 S1 You're standing on a cobblestone street surrounded by [...] [...] At 1 o'clock, about three steps away, a person is taking photos, which is a potential tripping hazard. [...] To move safely, proceed forward slightly to the left to avoid the person at 1 o'clock. S3 S3 Describe the situation to BLV people following GuideDog Standards.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59b44f7b-d7d8-4d6b-98df-6c19cb3bb38dCited by top-tier papers1
Ask how each one uses itBuilds on10
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
- Nuanced Perspectives Toward Disability Simulations from Digital Designers, Blind, Low Vision, and Color Blind PeopleGarreth W. TigwellCHI 2021 · 73 citations
- ORBIT: A Real-World Few-Shot Dataset for Teachable Object RecognitionDaniela Massiceti, Luisa M. Zintgraf, John Bronskill, Lida Theodorou et al.ICCV 2021 · 55 citations
Related papers
- WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian NavigationRafi Ibn Sultan, Hui Zhu, Xiangyu Zhou, Chengyin Li et al.CVPR 2026 · 4 citations
- VideoA11y: Method and Dataset for Accessible Video DescriptionChaoyu Li, Sid Padmanabhuni, Maryam S. Cheema, Hasti Seifi et al.CHI 2025 · 23 citations
- WalkVLM: Aid Visually Impaired People Walking by Vision Language ModelZhiqiang Yuan, Ting Zhang, Yeshuang Zhu, Jiapei Zhang et al.ICCV 2025 · 3 citations
- SceneScout: Towards AI-Driven Access to Street Level Imagery for Blind UsersGaurav Jain, Leah Findlater, Cole GleasonCHI 2026 · 2 citations
- Tiny Scales, Great Challenges: The Limits of Multimodal LLMs in Scale RecognitionJihang Jin, Ronghao Chen, Hao Zhang, Ziyan Liu et al.ACL 2026
