Science-T2I: Addressing Scientific Illusions in Image Synthesis
Jialuo Li, Wenhao Chai, Xingyu Fu, Haiyang Xu, Saining Xie
摘要
We present a novel approach to integrating scientific knowledge into generative models, enhancing their realism and consistency in image synthesis. First, we introduce Science-T2I, an expert-annotated adversarial dataset comprising adversarial 20k image pairs with 9k prompts, covering wide distinct scientific knowledge categories. Leveraging Science-T2I, we present SciScore, an end-toend reward model that refines the assessment of generated images based on scientific knowledge, which is achieved by augmenting both the scientific comprehension and visual capabilities of pre-trained CLIP model. Additionally, based on Science-T2I, we propose a two-stage training framework, comprising a supervised fine-tuning phase and a masked online fine-tuning phase, to incorporate scientific knowledge into existing generative models. Through comprehensive experiments, we demonstrate the effectiveness of our framework in establishing new standards for evaluating the scientific realism of generated content. Specifically, SciScore attains performance comparable to humanlevel, demonstrating a 5% improvement similar to evaluations conducted by experienced human evaluators. Furthermore, by applying our proposed fine-tuning method to FLUX, we achieve a performance enhancement exceeding 50% on SciScore.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive BenchmarkYang Shi, Yuhao Dong, Yue Ding, Yuran Wang 等CVPR 2026 · 被引用 35 次
- PAI-Bench: A Comprehensive Benchmark For Physical AIFengzhe Zhou, Jiannan Huang, Jialuo Li, Deva Ramanan 等CVPR 2026 · 被引用 32 次
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video GenerationHarold Haodong Chen, Haojian Huang, Qifeng Chen, Harry Yang 等NeurIPS 2025 · 被引用 25 次
- Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMsYujin Han, Hao Chen, Andi Han, Zhiheng Wang 等ICLR 2026 · 被引用 9 次
- PICABench: How Far are We from Physical Realistic Image Editing?Yuandong Pu, Le Zhuo, Songhao Han, Jinbo Xing 等ICLR 2026 · 被引用 7 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Style-GRPO: Semantic-Aware Preference Optimization for Image Style Transfer Guided by Reward ModelingJianbin Zhao, Chaoran Feng, Miao Yu, Yingtao Li 等CVPR 2026
- WISE: World Knowledge-Informed Semantic Evaluation for Text-to-Image GenerationYuwei Niu, Munan Ning, Mengren Zheng, Weiyang Jin 等ICML 2026 · 被引用 195 次
- Enhancing Reward Models for High-Quality Image Generation: Beyond Text-Image AlignmentYing Ba, Tianyu Zhang, Yalong Bai, Wenyi Mo 等ICCV 2025 · 被引用 13 次
- Aligning Synthetic Medical Images with Clinical Knowledge using Human FeedbackShenghuan Sun, Gregory M. Goldgof, Atul J. Butte, Ahmed M. AlaaNeurIPS 2023 · 被引用 27 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
