IC3: Image Captioning by Committee Consensus
David Chan, Austin Myers, Sudheendra Vijayanarasimhan, David A. Ross, John F. Canny
Abstract
If you ask a human to describe an image, they might do so in a thousand different ways. Image captioning models, on the other hand, are traditionally trained to generate a single "best" (most like a reference) caption. Unfortunately, doing so encourages captions that are informationally impoverished. Such captions often focus on only a subset of possible details, while ignoring other potentially useful information in the scene. In this work, we introduce a simple, yet novel, method: "Image Captioning by Committee Consensus" (IC 3 ), designed to generate a single caption that captures details from multiple viewpoints by sampling from the learned semantic space of a base captioning model, and carefully leveraging a large language model to synthesize these samples into a single comprehensive caption. Our evaluations show that humans rate captions produced by IC 3 more helpful than those produced by SOTA models more than two-thirds of the time, and IC 3 improves the performance of SOTA automated recall systems by up to 84%, outperforming single human-generated reference captions and indicating significant improvements over SOTA approaches for visual description. Code/Resources are available at https://davidmchan. github.io/caption-by-committee .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2d712de-4547-4947-bec5-6101d1cc211eCited by top-tier papers7
- Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation ModelsYuchen Hu, Chen Chen, Chao-Han Huck Yang, Chengwei Qin et al.NeurIPS 2024 · 14 citations
- "It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language ModelsKapil Garg, Xinru Tang, Jimin Heo, Dwayne R. Morgan et al.CHI 2026 · 2 citations
- SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language ModelsZhen Wan, Chao-Han Huck Yang, Yahan Yu, Jinchuan Tian et al.ACL 2025 · 2 citations
- Embodied Image Captioning: Self-Supervised Learning Agents for Spatially Coherent Image DescriptionsTommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio et al.ICCV 2025
- Visual Fact Checker: Enabling High-Fidelity Detailed Caption GenerationYunhao Ge, Xiaohui Zeng, Jacob Samuel Huffman, Tsung-Yi Lin et al.CVPR 2024
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
Related papers
- Show, Recall, and Tell: Image Captioning with Recall MechanismLi Wang, Zechen Bai, Yonghua Zhang, Hongtao LuAAAI 2020 · 73 citations
- Image Captioning with Multi-Context Synthetic DataFeipeng Ma, Yizhou Zhou, Fengyun Rao, Yueyi Zhang et al.AAAI 2024 · 22 citations
- InfoMetIC: An Informative Metric for Reference-free Image Caption EvaluationAnwen Hu, Shizhe Chen, Liang Zhang, Qin JinACL 2023 · 8 citations
- Multi-Perspective Video CaptioningYi Bin, Xindi Shang, Bo Peng, Yujuan Ding et al.ACM MM 2021 · 14 citations
- Learning Descriptive Image Captioning via Semipermeable Maximum Likelihood EstimationZihao Yue, Anwen Hu, Liang Zhang, Qin JinNeurIPS 2023 · 7 citations
