SlideAudit: A Dataset and Taxonomy for Automated Evaluation of Presentation Slides
Zhuohao Jerry Zhang, Ruiqi Chen, Mingyuan Zhong, Jacob O. Wobbrock
Abstract
Automated evaluation of specific graphic designs like presentation slides is an open problem. We present SlideAudit, a dataset for automated slide evaluation. We collaborated with design experts to develop a thorough taxonomy of slide design flaws. Our dataset comprises 2400 slides collected and synthesized from multiple sources, including a subset intentionally modified with specific design problems. We then fully annotated them using our taxonomy through strictly trained crowdsourcing from Prolific. To evaluate whether AI is capable of identifying design flaws, we compared multiple large language models under different prompting strategies, and with an existing design critique pipeline. We show that AI models struggle to accurately identify slide design flaws, with F1 scores ranging from 0.331 to 0.655. Notably, prompting techniques leveraging our taxonomy achieved the highest performance. We further conducted a remediation study to assess AI's potential for improving slides. Among 82.0% of slides that showed significant improvement, 87.8% of them were improved more with our taxonomy, further demonstrating its utility.
• Human-centered computing → Systems and tools for interaction design; Empirical studies in HCI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc144adc-9660-4eda-9b13-f46a5f38f4faCited by top-tier papers2
- VizCrit: Exploring Strategies for Displaying Computational Feedback in a Visual Design ToolMingyi Li, Mengyi Chen, Sarah Luo, Yining Cao et al.CHI 2026 · 1 citation
- As Content and Layout Co-Evolve: TangibleSite for Scaffolding Blind People's Webpage Design through Multimodal InteractionJiasheng Li, Zining Zhang, Zeyu Yan, Matthew Wong et al.CHI 2026 · 1 citation
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- SlideVQA: A Dataset for Document Visual Question Answering on Multiple ImagesRyota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa et al.AAAI 2023 · 178 citations
- Screen Recognition: Creating Accessibility Metadata for Mobile Applications from PixelsXiaoyi Zhang, Lilian de Greef, Amanda Swearngin, Samuel White et al.CHI 2021 · 145 citations
- UEyes: Understanding Visual Saliency across User Interface TypesYue Jiang, Luis A. Leiva, Hamed Rezazadegan Tavakoli, Paul R. B. Houssel et al.CHI 2023 · 100 citations
- Generating Automatic Feedback on UI Mockups with Large Language ModelsPeitong Duan, Jeremy Warner, Yang Li, Bjoern HartmannCHI 2024 · 81 citations
Related papers
- UICrit: Enhancing Automated Design Evaluation with a UI Critique DatasetPeitong Duan, Chin-Yi Cheng, Gang Li, Bjoern Hartmann et al.UIST 2024 · 22 citations
- 50 Shades of Deceptive Patterns: A Unified Taxonomy, Multimodal Detection, and Security ImplicationsZewei Shi, Ruoxi Sun, Jieshan Chen, Jiamou Sun et al.WWW 2025 · 12 citations
- LLMs Assist NLP Researchers: Critique Paper (Meta-)ReviewingJiangshu Du, Yibo Wang, Wenting Zhao, Zhongfen Deng et al.EMNLP 2024 · 14 citations
- Cognitive Biases in LLM-Assisted Software DevelopmentXinyi Zhou, Zeinadsadat Saghi, Sadra Sabouri, Rahul Pandita et al.ICSE 2026
- HILL: A Hallucination Identifier for Large Language ModelsFlorian Leiser, Sven Eckhardt, Valentin Leuthe, Merlin Knaeble et al.CHI 2024 · 67 citations
