Unblind your apps: predicting natural-language labels for mobile GUI components by deep learning
Jieshan Chen, Chunyang Chen, Zhenchang Xing, Xiwei Xu, Liming Zhu, Guoqiang Li, Jinshui Wang
Abstract
According to the World Health Organization(WHO), it is estimated that approximately 1.3 billion people live with some forms of vision impairment globally, of whom 36 million are blind. Due to their disability, engaging these minority into the society is a challenging problem. The recent rise of smart mobile phones provides a new solution by enabling blind users' convenient access to the information and service for understanding the world. Users with vision impairment can adopt the screen reader embedded in the mobile operating systems to read the content of each screen within the app, and use gestures to interact with the phone. However, the prerequisite of using screen readers is that developers have to add natural-language labels to the image-based components when they are developing the app. Unfortunately, more than 77% apps have issues of missing labels, according to our analysis of 10,408 Android apps. Most of these issues are caused by developers' lack of awareness and knowledge in considering the minority. And even if developers want to add the labels to UI components, they may not come up with concise and clear description as most of them are of no visual issues. To overcome these challenges, we develop a deep-learning based model, called LabelDroid, to automatically predict the labels of image-based buttons by learning from large-scale commercial apps in Google Play. The experimental results show that our model can make accurate predictions and the generated labels are of higher quality than that from real Android developers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fa2888ec-d95b-4f35-92cb-1fb9966cc3e8Cited by top-tier papers48
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- Screen Recognition: Creating Accessibility Metadata for Mobile Applications from PixelsXiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White et al.CHI 2021 · 145 citations
- Object detection for graphical user interface: old fashioned or deep learning or a combination?Jieshan Chen, Mulong Xie, Zhenchang Xing, Chunyang Chen et al.FSE 2020 · 144 citations
- Multi-Modal Repairs of Conversational Breakdowns in Task-Oriented DialogsToby Jia-Jun Li, Jingya Chen, Haijun Xia, Tom M. Mitchell et al.UIST 2020 · 98 citations
- Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of VarianceHung Viet Pham, Shangshu Qian, Jiannan Wang, Thibaud Lutellier et al.ASE 2020 · 91 citations
Builds on1
Related papers
- Data-driven accessibility repair revisited: on the effectiveness of generating labels for icons in Android appsForough Mehralian, Navid Salehnamadi, Sam MalekFSE 2021 · 47 citations
- Unblind Text Inputs: Predicting Hint-text of Text Input in Mobile Apps via LLMZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen et al.CHI 2024 · 29 citations
- AccessDroid: Detecting Screen Reader Accessibility Issues in Android Applications via Semantics TreesHan Zhou, Wei SongFSE 2026
- Accessibility issues in Android apps: state of affairs, sentiments, and ways forwardAbdulaziz Alshayban, Iftekhar Ahmed, Sam MalekICSE 2020 · 130 citations
- Bridging the Gap between Automated Intervention and Actual User Experience: A Mixed-Methods Study on Mobile Accessibility Issues for Screen Reader UsersSyed Fatiul Huq, Ziyao He, Yirui He, Sam MalekCHI 2026 · 1 citation
