FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
Rongyao Fang, Aldrich Yu, Chengqi Duan, Linjiang Huang, Shuai Bai, Yuxuan Cai, Kun Wang, Si Liu, Xihui Liu, Hongsheng Li
摘要
The advancement of open-source text-to-image (T2I) models has been hindered by the absence of large-scale, reasoning-focused datasets and comprehensive evaluation benchmarks, resulting in a performance gap compared to leading closed-source systems. To address this challenge, We introduce FLUX-Reason-6M and PRISM-Bench (Precise and Robust Image Synthesis Measurement Benchmark). FLUX-Reason-6M is a massive dataset consisting of 6 million high-quality FLUX-generated images and 20 million bilingual (English and Chinese) descriptions specifically designed to teach complex reasoning. The image are organized according to six key characteristics: Imagination, Entity, Text rendering, Style, Affection, and Composition, and design explicit Generation Chain-of-Thought (GCoT) to provide detailed breakdowns of image generation steps. The whole data curation takes 15,000 A100 GPU days, providing the community with a resource previously unattainable outside of large industrial labs. PRISM-Bench offers a novel evaluation standard with seven distinct tracks, including a formidable Long Text challenge using GCoT. Through carefully designed prompts, it utilizes advanced visionlanguage models for nuanced human-aligned assessment of prompt-image alignment and image aesthetics. Our extensive evaluation of 19 leading models on PRISM-Bench reveals critical performance gaps and highlights specific areas requiring improvement. Our dataset, benchmark, and evaluation code are released to catalyze the next wave of reasoning-oriented T2I generation. Composition: Cute Morkie puppy perched inside a pink box labeled "Fabulous". Adorned with a matching pink bow and surrounded by fluffy feathers and vibrant pink flowers, the puppy exudes charm and playfulness. Affection: Enveloped in a world of soft textures and vibrant hues, this Morkie puppy embodies the essence of innocence and wonder, radiating an irresistible charm that ignites a heartwarming sense of joy and hope. GCoT: Cute and charming, a Morkie puppy sits adorably perched inside a pastel pink cardboard box labeled "Fabulous" in bold black script. The puppy, with its fluffy white fur accented by brown patches around its ears and face, exudes an irresistible charm, enhanced by a matching pink satin bow tied neatly atop its head. Its small paws, tipped with dark pads, rest gently on the edge of the box, adding a playful touch to the composition. Surrounding the puppy are fluffy white feathers and vibrant pink flowers, creating a whimsical and enchanting atmosphere. The background is a soft, uniform pink, which complements the overall color scheme and draws attention to the central subject. The image captures a perfect blend of innocence and playfulness, evoking a heartwarming sense of joy and hope through its soft textures, vibrant hues, and meticulous arrangement.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?Ouxiang Li, Yuan Wang, Xinting Hu, Huijuan Huang 等ICLR 2026 · 被引用 39 次
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow MatchingRun Luo, Xiaobo Xia, Lu Wang, Longze Chen 等ICLR 2026 · 被引用 22 次
- Factuality Matters: When Image Generation and Editing Meet Structured VisualsLe Zhuo, Songhao Han, Yuandong Pu, Boxiang Qiu 等ICLR 2026 · 被引用 15 次
- TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text RenderingHanshen Zhu, Yuliang Liu, Xuecheng Wu, An-Lan Wang 等CVPR 2026 · 被引用 15 次
- ThinkGen: Generalized Thinking for Visual GenerationSiyu Jiao, Yiheng Lin, Yujie Zhong, Qi She 等CVPR 2026 · 被引用 12 次
它引用的顶会 Paper24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- GIR-Bench: Versatile Benchmark for Generating Images with ReasoningHongxiang Li, Yaowei Li, Bin Lin, Yuwei Niu 等ICLR 2026 · 被引用 15 次
- MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine TranslationGengluo Li, Chengquan Zhang, Yupu Liang, Huawen Shen 等CVPR 2026 · 被引用 6 次
- R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image GenerationKaijie Chen, Zihao Lin, Zhiyang Xu, Ying Shen 等EMNLP 2025
- Will Multimodal Models Be Dazzled by Multi-Image Visual Puzzles?zhi zhu, YaoQi Fan, Zhe Chen, Yue Cao 等CVPR 2026
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
