Context Matters for Image Descriptions for Accessibility: Challenges for Referenceless Evaluation Metrics
Elisa Kreiss, Cynthia L. Bennett, Shayan Hooshmand, Eric Zelikman, Meredith Ringel Morris, Christopher Potts
Abstract
Few images on the Web receive alt-text descriptions that would make them accessible to blind and low vision (BLV) users. Image-based NLG systems have progressed to the point where they can begin to address this persistent societal problem, but these systems will not be fully successful unless we evaluate them on metrics that guide their development correctly. Here, we argue against current referenceless metrics – those that don't rely on human-generated ground-truth descriptions – on the grounds that they do not align with the needs of BLV users. The fundamental shortcoming of these metrics is that they do not take context into account, whereas contextual information is highly valued by BLV users. To substantiate these claims, we present a study with BLV participants who rated descriptions along a variety of dimensions. An in-depth analysis reveals that the lack of context-awareness makes current referenceless metrics inadequate for advancing image accessibility. As a proof-of-concept, we provide a contextual version of the referenceless metric CLIPScore which begins to address the disconnect to the BLV data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- WorldScribe: Towards Context-Aware Live Visual DescriptionsRuei-Che Chang, Yuxuan Liu, Anhong GuoUIST 2024 · 54 citations
- Investigating Use Cases of AI-Powered Scene Description Applications for Blind and Low Vision PeopleRicardo E. Gonzalez Penuela, Jazmin Collins, Cynthia L. Bennett, Shiri AzenkotCHI 2024 · 44 citations
- Unblind Text Inputs: Predicting Hint-text of Text Input in Mobile Apps via LLMZhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen et al.CHI 2024 · 29 citations
- Alt-Text with Context: Improving Accessibility for Images on TwitterNikita Srivatsan, Sofía Samaniego, Omar Florez, Taylor Berg-KirkpatrickICLR 2024 · 9 citations
- ContextRef: Evaluating Referenceless Metrics for Image Description GenerationElisa Kreiss, Eric Zelikman, Christopher Potts, Nick HaberICLR 2024 · 6 citations
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- "Person, Shoes, Tree. Is the Person Naked?" What People with Vision Impairments Want in Image DescriptionsAbigale Stangl, Meredith Ringel Morris, Danna GurariCHI 2020 · 136 citations
- Twitter A11y: A Browser Extension to Make Twitter Images AccessibleCole Gleason, Amy Pavel, Emma McCamey, Christina Low et al.CHI 2020 · 123 citations
Related papers
- HICEScore: A Hierarchical Metric for Image Captioning EvaluationZequn Zeng, Jianqiao Sun, Hao Zhang, Tiansheng Wen et al.ACM MM 2024 · 3 citations
- LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis EvaluationYujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang et al.NeurIPS 2023 · 119 citations
- SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption EvaluationXiaofu Chen, Israfel Salazar, Yova KementchedjhievaEMNLP 2025
- Language Model Augmented Relevance ScoreRuibo Liu, Jason Wei, Soroush VosoughiACL 2021
- Prompt Expansion for Adaptive Text-to-Image GenerationSiddhartha Datta, Alexander Ku, Deepak Ramachandran, Peter AndersonACL 2024 · 3 citations
