AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
Vatsal Malaviya, Agneet Chatterjee, Maitreya Patel, Yezhou Yang, Chitta Baral
Abstract
Text-to-Image (T2I) models have recently achieved remarkable success in generating images from textual descriptions. However, challenges still persist in accurately rendering complex scenes where actions and interactions form the primary semantic focus. Our key observation in this work is that T2I models frequently struggle to capture nuanced and often implicit attributes inherent in action depiction, leading to generating images that lack key contextual details. To enable systematic evaluation, we introduce AcT2I, a benchmark designed to evaluate the performance of T2I models in generating images from action-centric prompts. We experimentally validate that leading T2I models do not fare well on AcT2I. We further hypothesize that this shortcoming arises from the incomplete representation of the inherent attributes and contextual dependencies in the training corpora of existing T2I models. We build upon this by developing a trainingfree, knowledge distillation technique utilizing Large Language Models to address this limitation. Specifically, we enhance prompts by incorporating dense information across three dimensions, observing that injecting prompts with temporal details significantly improves image generation accuracy, with our best model achieving an increase of 72%. Our findings highlight the limitations of current T2I methods in generating images that require complex reasoning and demonstrate that integrating linguistic knowledge in a systematic way can notably advance the generation of nuanced and contextually accurate images. Project
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a859fef2-a0f6-4982-b9a4-4f5e934ca027Builds on10
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani et al.NeurIPS 2023 · 462 citations
- InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image GenerationXingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng et al.ICLR 2024 · 358 citations
- Animal Kingdom: A Large and Diverse Dataset for Animal Behavior UnderstandingXun Long Ng, Kian Eng Ong, Qichen Zheng, Yun Ni et al.CVPR 2022 · 102 citations
Related papers
- TF-TI2I: Training-Free Text-And-Image-To-Image Generation via Multi-Modal Implicit-Context Learning in Text-To-Image ModelsTeng-Fang Hsiao, Bo-Kai Ruan, Yi-Lun Wu, Tzu-Ling Lin et al.ICCV 2025 · 3 citations
- Color Me Correctly: Bridging Perceptual Color Spaces and Text Embeddings for Improved Diffusion GenerationSung-Lin Tsai, Bo-Lun Huang, Yu-Ting Shen, Cheng-Yu Yeo et al.ACM MM 2025 · 3 citations
- DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?Qirui Jiao, Daoyuan Chen, Yilun Huang, Xika Lin et al.ICML 2026
- Long-Text-to-Image Generation via Compositional Prompt DecompositionJen-Yuan Huang, Tong Lin, Yilun DuICLR 2026 · 1 citation
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?Ouxiang Li, Yuan Wang, Xinting Hu, Huijuan Huang et al.ICLR 2026 · 39 citations
