Lune

CHI2026Top-tier venue

Do Entropic Measurements of the Diversity of AI-generated Images Match Human Judgement?

Kazjon Grace, Francisco Javier Ibarrola, Jody Watts, Shu Takahashi, Parth Bhargava, Eduardo Velloso

2026Year
1Citations

Abstract

This paper proposes that the ability to generate diverse outputs in response to a single prompt is necessary for text-to-image models to become more effective creativity support tools. It formalises the problem of measuring the diversity of generated text and images, with an emphasis on interactive, exploratory use in open-ended and creative tasks. It suggests, motivated by research in the psychology of creativity, that diversity should sit alongside image quality and fit-to-prompt as critical measures in this setting. The paper adapts several diversity measures from the literature to this task, then explores how they compare to human diversity ratings. These evaluations show that algorithmic measures of diversity can be a useful proxy for human ratings, with both declining in accuracy as the difficulty of the task increases. The paper concludes with an exploratory qualitative analysis of the factors involved in human diversity judgments to guide future research in this emerging area.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 21e3dd98-4319-4789-ad99-a507e8a7cb78

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines