Screen Recognition: Creating Accessibility Metadata for Mobile Applications from Pixels
Xiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White, Kyle I. Murray, Lisa Yu, Qi Shan, Jeffrey Nichols, Jason Wu, Chris Fleizach, Aaron Everitt, Jeffrey P. Bigham
摘要
Many accessibility features available on mobile platforms require applications (apps) to provide complete and accurate metadata describing user interface (UI) components. Unfortunately, many apps do not provide sufficient metadata for accessibility features to work as expected. In this paper, we explore inferring accessibility metadata for mobile apps from their pixels, as the visual interfaces often best reflect an app’s full functionality. We trained a robust, fast, memory-efficient, on-device model to detect UI elements using a dataset of 77,637 screens (from 4,068 iPhone apps) that we collected and annotated. To further improve UI detections and add semantic information, we introduced heuristics (e.g., UI grouping and ordering) and additional models (e.g., recognize UI content, state, interactivity). We built Screen Recognition to generate accessibility metadata to augment iOS VoiceOver. In a study with 9 screen reader users, we validated that our approach improves the accessibility of existing mobile apps, enabling even previously inaccessible apps to be used.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper57
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu 等ICML 2023 · 被引用 426 次
- Enabling Conversational Interaction with Mobile UI using Large Language ModelsBryan Wang, Gang Li, Yang LiCHI 2023 · 被引用 149 次
- Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningBryan Wang, Gang Li, Xin Zhou, Zhourong Chen 等UIST 2021 · 被引用 97 次
- AutoDroid: LLM-powered Task Automation in AndroidHao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao 等MobiCom 2024 · 被引用 94 次
- Make LLM a Testing Expert: Bringing Human-like Interaction to Mobile GUI Testing via Functionality-aware DecisionsZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen 等ICSE 2024 · 被引用 81 次
它引用的顶会 Paper5
- Object detection for graphical user interface: old fashioned or deep learning or a combination?Jieshan Chen, Mulong Xie, Zhenchang Xing, Chunyang Chen 等FSE 2020 · 被引用 144 次
- Twitter A11y: A Browser Extension to Make Twitter Images AccessibleCole Gleason, Amy Pavel, Emma McCamey, Christina Low 等CHI 2020 · 被引用 123 次
- Unblind your apps: predicting natural-language labels for mobile GUI components by deep learningJieshan Chen, Chunyang Chen, Zhenchang Xing, Xiwei Xu 等ICSE 2020 · 被引用 101 次
- DeepIntent: Deep Icon-Behavior Learning for Detecting Intention-Behavior Discrepancy in Mobile AppsShengqu Xi, Shao Yang, Xusheng Xiao, Yuan Yao 等CCS 2019 · 被引用 74 次
- From Lost to Found: Discover Missing UI Design Semantics through Recovering Missing TagsChunyang Chen, Sidong Feng, Zhengyang Liu, Zhenchang Xing 等CSCW 2020 · 被引用 41 次
相关 Paper
- WebUI: A Dataset for Enhancing Visual UI Understanding with Web SemanticsJason Wu, Siyan Wang, Siman Shen, Yi-Hao Peng 等CHI 2023 · 被引用 49 次
- Screen Parsing: Towards Reverse Engineering of UI Models from ScreenshotsJason Wu, Xiaoyi Zhang, Jeffrey Nichols, Jeffrey P. BighamUIST 2021 · 被引用 62 次
- Never-ending Learning of User InterfacesJason Wu, Rebecca Krosnick, Eldon Schoop, Amanda Swearngin 等UIST 2023 · 被引用 17 次
- Towards Complete Icon Labeling in Mobile ApplicationsJieshan Chen, Amanda Swearngin, Jason Wu, Titus Barik 等CHI 2022 · 被引用 30 次
- ScreenAudit: Detecting Screen Reader Accessibility Errors in Mobile Apps Using Large Language ModelsMingyuan Zhong, Ruolin Chen, Xia Chen, James Fogarty 等CHI 2025 · 被引用 15 次
