ACL2026
Evaluating Visual Narrative Coherence in Story Visualization via Diversified Storylines
Minha Jhang, Kyeongman Park, Hyukhun Koh, Kyomin Jung
Abstract
Story visualization requires generating a coherent sequence of images that collectively form a narrative, yet existing evaluation metrics and datasets often overlook visual continuity and narrative diversity. In this paper, we introduce the Visual Context-Aware Metric for Story Visualization (VCMS) , which uses large vision-language models to jointly assess caption fidelity and inter-image consistency, achieving Spearman’s correlation comparable to human agreement on two benchmarks. Also, to address the shortcomings of narrowly defined datasets with low diversity, we propose a diffusion-augmented evaluation pipeline that blends diverse and controlled narrative elements at adjustable ratios, enabling the creation of evaluation sets tailored to specific objectives. By combining VCMS with this pipeline, we provide a scalable, human-aligned framework for evaluating story visualization models.