Alt-Text with Context: Improving Accessibility for Images on Twitter
Nikita Srivatsan, Sofía Samaniego, Omar Florez, Taylor Berg-Kirkpatrick
Abstract
In this work we present an approach for generating alternative text (or alt-text) descriptions for images shared on social media, specifically Twitter. More than just a special case of image captioning, alt-text is both more literally descriptive and context-specific. Also critically, images posted to Twitter are often accompanied by user-written text that despite not necessarily describing the image may provide useful context that if properly leveraged can be informative. We address this task with a multimodal model that conditions on both textual information from the associated social media post as well as visual signal from the image, and demonstrate that the utility of these two information sources stacks. We put forward a new dataset of 371k images paired with alt-text and tweets scraped from Twitter and evaluate on it across a variety of automated metrics as well as human evaluation. We show that our approach of conditioning on both tweet text and visual information significantly outperforms prior work, by more than 2x on BLEU@4.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 472942c4-5018-460e-ad72-33576b2e42f4Cited by top-tier papers2
- Influencer: Empowering Everyday Users in Creating Promotional Posts via AI-infused Exploration and CustomizationXuye Liu, Annie Sun, Pengcheng An, Tengfei Ma et al.CHI 2025 · 8 citations
- MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text GenerationJinlan Fu, Shenzhen Huangfu, Hao Fei, Yichong Huang et al.ACM MM 2025
Builds on8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Twitter A11y: A Browser Extension to Make Twitter Images AccessibleCole Gleason, Amy Pavel, Emma McCamey, Christina Low et al.CHI 2020 · 123 citations
Related papers
- Widget Captioning: Generating Natural Language Description for Mobile User Interface ElementsYang Li, Gang Li, Luheng He, Jingjie Zheng et al.EMNLP 2020 · 46 citations
- Exploiting BERT for Multimodal Target Sentiment Classification through Input Space TranslationZaid Khan, Yun FuACM MM 2021 · 192 citations
- Cross-Media Keyphrase Prediction: A Unified Framework with Multi-Modality Multi-Head Attention and Image WordingsYue Wang, Jing Li, Michael R. Lyu, Irwin KingEMNLP 2020 · 11 citations
- Multimodal Relation Extraction with Efficient Graph AlignmentChangmeng Zheng, Junhao Feng, Ze Fu, Yi Cai et al.ACM MM 2021 · 134 citations
- CapOnImage: Context-driven Dense-Captioning on ImageYiqi Gao, Xinglin Hou, Yuanmeng Zhang, Tiezheng Ge et al.EMNLP 2022 · 5 citations
