ImageInWords: Unlocking Hyper-Detailed Image Descriptions
Roopal Garg, Andrea Burns, Burcu Karagol Ayan, Yonatan Bitton, Ceslee Montgomery, Yasumasa Onoe, Andrew Bunner, Ranjay Krishna, Jason Baldridge, Radu Soricut
Abstract
Despite the longstanding adage "an image is worth a thousand words," generating accurate hyper-detailed image descriptions remains unsolved. Trained on short web-scraped imagetext, vision-language models often generate incomplete descriptions with visual inconsistencies. We address this via a novel data-centric approach with ImageInWords (IIW), a carefully designed human-in-the-loop framework for curating hyper-detailed image descriptions. Human evaluations on IIW data show major gains compared to recent datasets (+66%) and GPT-4V (+48%) across comprehensiveness, specificity, hallucinations, and more. We also show that fine-tuning with IIW data improves these metrics by +31% against models trained with prior work, even with only 9k samples. Lastly, we evaluate IIW models with text-to-image generation and vision-language reasoning tasks. Our generated descriptions result in the highest fidelity images, and boost compositional reasoning by up to 6% on ARO, SVO-Probes, and Winoground datasets. We release the IIW-Eval benchmark with human judgement labels, object and image-level annotations from our framework, and existing image caption datasets enriched via IIW-model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers26
- Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed PerceptionZiyang Ma, Ruiyang Xu, Zhenghao Xing, Yunfei Chu et al.ICLR 2026 · 36 citations
- LoTLIP: Improving Language-Image Pre-training for Long Text UnderstandingWei Wu, Kecheng Zheng, Shuailei Ma, Fan Lu et al.NeurIPS 2024 · 35 citations
- Cycle Consistency as Reward: Learning Image-Text Alignment Without Human PreferencesHyojin Bahng, Caroline Chan, Frédo Durand, Phillip IsolaICCV 2025 · 25 citations
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-FollowingTianyi Xiong, Yi Ge, Ming Li, Zuolong Zhang et al.CVPR 2026 · 16 citations
- ScaleCap: Scalable Image Captioning via Dual-Modality DebiasingLong Xing, Qidong Huang, Xiaoyi Dong, Pan Zhang et al.ICLR 2026 · 11 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
Related papers
- Caption This, Reason That: VLMs Caught in the MiddleZihan Weng, Lucas Gomez, Taylor W. Webb, Pouya BashivanNeurIPS 2025 · 3 citations
- IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMsDavid Ma, Yuanxing Zhang, Jincheng Ren, Jiawei Guo et al.ICLR 2026 · 5 citations
- Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language ModelsJeonghwan Kim, Heng JiEMNLP 2024 · 4 citations
- Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal LearningTianyi Bai, Yuxuan Fan, Jiantao Qiu, Fupeng Sun et al.NeurIPS 2025 · 12 citations
- Probing Visual Language Priors in VLMsTiange Luo, Ang Cao, Gunhee Lee, Justin Johnson et al.ICML 2025
