Alt-Text with Context: Improving Accessibility for Images on Twitter
Nikita Srivatsan, Sofía Samaniego, Omar Florez, Taylor Berg-Kirkpatrick
摘要
In this work we present an approach for generating alternative text (or alt-text) descriptions for images shared on social media, specifically Twitter. More than just a special case of image captioning, alt-text is both more literally descriptive and context-specific. Also critically, images posted to Twitter are often accompanied by user-written text that despite not necessarily describing the image may provide useful context that if properly leveraged can be informative. We address this task with a multimodal model that conditions on both textual information from the associated social media post as well as visual signal from the image, and demonstrate that the utility of these two information sources stacks. We put forward a new dataset of 371k images paired with alt-text and tweets scraped from Twitter and evaluate on it across a variety of automated metrics as well as human evaluation. We show that our approach of conditioning on both tweet text and visual information significantly outperforms prior work, by more than 2x on BLEU@4.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Influencer: Empowering Everyday Users in Creating Promotional Posts via AI-infused Exploration and CustomizationXuye Liu, Annie Sun, Pengcheng An, Tengfei Ma 等CHI 2025 · 被引用 8 次
- MCM-DPO: Multifaceted Cross-Modal Direct Preference Optimization for Alt-text GenerationJinlan Fu, Shenzhen Huangfu, Hao Fei, Yichong Huang 等ACM MM 2025
它引用的顶会 Paper8
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Twitter A11y: A Browser Extension to Make Twitter Images AccessibleCole Gleason, Amy Pavel, Emma McCamey, Christina Low 等CHI 2020 · 被引用 123 次
相关 Paper
- Widget Captioning: Generating Natural Language Description for Mobile User Interface ElementsYang Li, Gang Li, Luheng He, Jingjie Zheng 等EMNLP 2020 · 被引用 46 次
- Exploiting BERT for Multimodal Target Sentiment Classification through Input Space TranslationZaid Khan, Yun FuACM MM 2021 · 被引用 192 次
- Cross-Media Keyphrase Prediction: A Unified Framework with Multi-Modality Multi-Head Attention and Image WordingsYue Wang, Jing Li, Michael R. Lyu, Irwin KingEMNLP 2020 · 被引用 11 次
- Multimodal Relation Extraction with Efficient Graph AlignmentChangmeng Zheng, Junhao Feng, Ze Fu, Yi Cai 等ACM MM 2021 · 被引用 134 次
- CapOnImage: Context-driven Dense-Captioning on ImageYiqi Gao, Xinglin Hou, Yuanmeng Zhang, Tiezheng Ge 等EMNLP 2022 · 被引用 5 次
