DIMCIM: A Quantitative Evaluation Framework for Default-Mode Diversity and Generalization in Text-to-Image Generative Models
Revant Teotia, Candace Ross, Karen Ullrich, Sumit Chopra, Adriana Romero-Soriano, Melissa Hall, Matthew J. Muckley
摘要
Recent advances in text-to-image (T2I) models have achieved impressive quality and consistency. However, this has come at the cost of representation diversity. While automatic evaluation methods exist for benchmarking model diversity, they either require reference image datasets or lack specificity about the kind of diversity measured, limiting their adaptability and interpretability. To address this gap, we introduce the Does-it/Can-it framework, DIM-CIM, a reference-free measurement of default-mode diversity ("Does" the model generate images with expected attributes?) and generalization capacity ("Can" the model generate diverse attributes for a particular concept?). We construct the COCO-DIMCIM benchmark, which is seeded with COCO concepts and captions and augmented by a large language model. With COCO-DIMCIM, we find that widely-used models improve in generalization at the cost of default-mode diversity when scaling from 1.5B to 8.1B parameters. DIMCIM also identifies fine-grained failure cases, such as attributes that are generated with generic prompts but are rarely generated when explicitly requested. Finally, we use DIMCIM to evaluate the training data of a T2I model and observe a correlation of 0.85 between diversity in training images and default-mode diversity. Our work provides a flexible and interpretable framework for assessing T2I model diversity and generalization, enabling a more comprehensive understanding of model performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Taming Preference Mode Collapse via Directional Decoupling Alignment in Diffusion Reinforcement LearningChubin Chen, Sujie Hu, Jiashu Zhu, Meiqi Wu 等CVPR 2026 · 被引用 28 次
- GeoDiv: Framework for Measuring Geographical Diversity in Text-to-Image ModelsAbhipsa Basu, Mohana Singh, Shashank Agnihotri, Margret Keuper 等ICLR 2026 · 被引用 3 次
- On-the-fly Repulsion in the Contextual Space for Rich Diversity in Diffusion TransformersOmer Dahary, Benaya Koren, Daniel Garibi, Daniel Cohen-OrSIGGRAPH 2026 · 被引用 1 次
- Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow MatchingJingxuan Wu, Zhenglin Wan, Xingrui Yu, Yuzhe YANG 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper16
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
相关 Paper
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- The Intricate Dance of Prompt Complexity, Quality, Diversity and Consistency in T2I ModelsXiaofeng Zhang, Aaron C. Courville, Michal Drozdzal, Adriana Romero-SorianoICLR 2026 · 被引用 6 次
- ScImage: How good are multimodal large language models at scientific text-to-image generation?Leixin Zhang, Steffen Eger, Yinjie Cheng, Weihe Zhai 等ICLR 2025
- Trade-Offs in Image Generation: How Do Different Dimensions Interact?Sicheng Zhang, Binzhu Xie, Zhonghao Yan, Yuli Zhang 等ICCV 2025 · 被引用 1 次
- Do Entropic Measurements of the Diversity of AI-generated Images Match Human Judgement?Kazjon Grace, Francisco Javier Ibarrola, Jody Watts, Shu Takahashi 等CHI 2026 · 被引用 1 次
