Does GenAI Make Usability Testing Obsolete?
Ali Ebrahimi Pourasad, Walid Maalej
Abstract
Ensuring usability is crucial for the success of mobile apps. Usability issues can compromise user experience and negatively impact the perceived app quality. This paper presents UX-LLM, a novel tool powered by a Large Vision-Language Model that predicts usability issues in iOS apps. To evaluate the performance of UX-LLM, we predicted usability issues in two open-source apps of a medium complexity and asked two usability experts to assess the predictions. We also performed traditional usability testing and expert review for both apps and compared the results to those of UX-LLM. UX-LLM demonstrated precision ranging from 0.61 and 0.66 and recall between 0.35 and 0.38, indicating its ability to identify valid usability issues, yet failing to capture the majority of issues. Finally, we conducted a focus group with an app development team of a capstone project developing a transit app for visually impaired persons. The focus group expressed positive perceptions of UX-LLM as it identified unknown usability issues in their app. However, they also raised concerns about its integration into the development workflow, suggesting potential improvements. Our results show that UX-LLM cannot fully replace traditional usability evaluation methods but serves as a valuable supplement particularly for small teams with limited resources, to identify issues in less common user paths, due to its ability to inspect the source code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- ViBR: Automated Bug Replay from Video-Based Reports using Vision-Language ModelsSidong Feng, Dingbang Wang, Nikola Tomic, Tingting Yu et al.FSE 2026 · 1 citation
- Recommending Usability Improvements with Multimodal Large Language ModelsSebastian Lubos, Alexander Felfernig, Damian Garber, Viet-Man Le et al.FSE 2026
- LikeThis! Empowering App Users to Submit UI Improvement Suggestions Instead of ComplaintsJialiang Wei, Ali Ebrahimi Pourasad, Walid MaalejICSE 2026
Builds on5
- The Unreliability of Explanations in Few-shot Prompting for Textual ReasoningXi Ye, Greg DurrettNeurIPS 2022 · 272 citations
- Enhancing UX Evaluation Through Collaboration with Conversational AI Assistants: Effects of Proactive Dialogue and TimingEmily Kuang, Minghao Li, Mingming Fan, Kristen ShinoharaCHI 2024 · 47 citations
- Semantic Web Accessibility Testing via Hierarchical Visual AnalysisMohammad Bajammal, Ali MesbahICSE 2021 · 29 citations
- AccessiText: automated detection of text accessibility issues in Android appsAbdulaziz Alshayban, Sam MalekFSE 2022 · 28 citations
- LayoutDM: Transformer-based Diffusion Model for Layout GenerationShang Chai, Liansheng Zhuang, Fengying YanCVPR 2023
Related papers
- Peeking Ahead of the Field Study: Exploring VLM Personas as Support Tools for Embodied Studies in HCIXinyue Gui, Ding Xia, Mark Colley, Yuan Li et al.CHI 2026 · 1 citation
- AXNav: Replaying Accessibility Tests from Natural LanguageMaryam Taeb, Amanda Swearngin, Eldon Schoop, Ruijia Cheng et al.CHI 2024 · 51 citations
- Understanding the Use of a Large Language Model-Powered Guide to Make Virtual Reality Accessible for Blind and Low Vision PeopleJazmin Collins, Sharon Y. Lin, Tianqi Liu, Andrea Stevenson Won et al.CHI 2026 · 3 citations
- Evaluating Object Hallucination in Large Vision-Language ModelsYifan Li, Yifan Du, Kun Zhou, Jinpeng Wang et al.EMNLP 2023 · 344 citations
- "It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language ModelsKapil Garg, Xinru Tang, Jimin Heo, Dwayne R. Morgan et al.CHI 2026 · 2 citations
