Towards LLM-powered Assistive Drone for Blind and Low Vision Users
Yize Wei, Ibnu Taimiyyah Bin Adam, Hanjun Wu, Moritz Messerschmidt, Wei Tsang Ooi, Christophe Jouffrais, Suranga Nanayakkara
Abstract
Drones have gained traction as a versatile form of assistive robots for Blind and Low Vision (BLV) people. Nonetheless, novel interaction techniques are required to enable BLV people to communicate with drones naturally. In this work, we built an LLM-powered assistive drone for BLV users. We leverage an LLM to translate high-level user goals to step-by-step instructions for the drone and to extract visual information from the images. Through a formative study with BLV users (N=9), we identified envisioned use cases and desired interaction modalities. Then, we took a participatory and iterative approach to build a prototype, incorporating feedback received from 3 BLV users, as well as 5 domain experts. Finally, we conducted a user study with an additional 6 BLV participants to evaluate the iterated prototype, and received positive feedback. This work is contributing to a growing body of research on harnessing the power of LLMs to build a more inclusive world.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8034d617-5a4b-42c8-87fd-0e9449726cf5Builds on26
- Re-examining Whether, Why, and How Human-AI Interaction Is Uniquely Difficult to DesignQian Yang, Aaron Steinfeld, Carolyn P. Rosé, John ZimmermanCHI 2020 · 604 citations
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet et al.NeurIPS 2024 · 166 citations
- Bridging the Gulf of Envisioning: Cognitive Challenges in Prompt Based Interactions with LLMsHariharan Subramonyam, Roy Pea, Christopher Lawrence Pondoc, Maneesh Agrawala et al.CHI 2024 · 137 citations
- Robots for Inclusive Play: Co-designing an Educational Game With Visually Impaired and sighted ChildrenOussama Metatla, Sandra Bardot, Clare Cullen, Marcos Serrano et al.CHI 2020 · 111 citations
- Above and Beyond: A Scoping Review of Domains and Applications for Human-Drone InteractionViviane Herdel, Lee J. Yamin, Jessica R. CauchardCHI 2022 · 66 citations
Related papers
- Understanding the Use of a Large Language Model-Powered Guide to Make Virtual Reality Accessible for Blind and Low Vision PeopleJazmin Collins, Sharon Y. Lin, Tianqi Liu, Andrea Stevenson Won et al.CHI 2026 · 3 citations
- How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision PeopleRicardo E. Gonzalez Penuela, Crescentia Jung, Sharon Y. Lin, Ruiying Hu et al.CHI 2026 · 1 citation
- SceneScout: Towards AI-Driven Access to Street Level Imagery for Blind UsersGaurav Jain, Leah Findlater, Cole GleasonCHI 2026 · 2 citations
- Not Seeing the Whole Picture: Challenges and Opportunities in Using AI for Co-Making Physical, DIY-AT for People with Visual ImpairmentsBen Kosa, Hsuanling Lee, Jasmine Li, Sanbrita Mondal et al.CHI 2026 · 1 citation
- DR-VQA: Decompose-then-Reconstruct for Visual Question Answering in BLV AssistanceBocheng Pan, Hailong Shi, Xingyu GaoACM MM 2025
