ACL2026
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
Junhyeok Kim, Jaewoo Park, Junhee Park, Sangeyl Lee, Jiwan Chung, Jisung Kim, Ji Hoon Joung, Youngjae Yu
被引用 7 次
摘要
For people affected by blindness and low vision (BLV), safe and independent navigation remains a major challenge, impacting over 2.2 billion individuals worldwide. Although multimodal large language models (MLLMs) offer new opportunities for assistive navigation, progress has been limited by the scarcity of accessibility-aware datasets, because creating them requires labor-intensive expert annotation. To this end, we introduce GUIDEDOG, a novel dataset containing 22K image-description pairs (2K human-verified) capturing real-world pedestrian scenes across 46 countries. Our human-AI pipeline shifts annotation from generation to verification, grounded in established BLV guidance standards from experts and research, improving scalability while maintaining quality. We also present GUIDEDOGQA, an 818-sample benchmark evaluating object recognition and depth perception. Experiments reveal that depth perception and adherence to these standards remain challenging for current MLLMs. * Equal Contribution. † Work done while at SK Telecom. S2 S2 S1 S1 You're standing on a cobblestone street surrounded by [...] [...] At 1 o'clock, about three steps away, a person is taking photos, which is a potential tripping hazard. [...] To move safely, proceed forward slightly to the left to avoid the person at 1 o'clock. S3 S3 Describe the situation to BLV people following GuideDog Standards.