Lune

EMNLP2023Top-tier venue

IC3: Image Captioning by Committee Consensus

David Chan, Austin Myers, Sudheendra Vijayanarasimhan, David A. Ross, John F. Canny

2023Year
8Citations
7Top-tier citations

Abstract

If you ask a human to describe an image, they might do so in a thousand different ways. Image captioning models, on the other hand, are traditionally trained to generate a single "best" (most like a reference) caption. Unfortunately, doing so encourages captions that are informationally impoverished. Such captions often focus on only a subset of possible details, while ignoring other potentially useful information in the scene. In this work, we introduce a simple, yet novel, method: "Image Captioning by Committee Consensus" (IC 3 ), designed to generate a single caption that captures details from multiple viewpoints by sampling from the learned semantic space of a base captioning model, and carefully leveraging a large language model to synthesize these samples into a single comprehensive caption. Our evaluations show that humans rate captions produced by IC 3 more helpful than those produced by SOTA models more than two-thirds of the time, and IC 3 improves the performance of SOTA automated recall systems by up to 84%, outperforming single human-generated reference captions and indicating significant improvements over SOTA approaches for visual description. Code/Resources are available at https://davidmchan. github.io/caption-by-committee .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f2d712de-4547-4947-bec5-6101d1cc211e

Cited by top-tier papers7

Ask how each one uses it

Builds on16

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines