Image Content Generation with Causal Reasoning
Xiaochuan Li, Baoyu Fan, Runze Zhang, Liang Jin, Di Wang, Zhenhua Guo, Yaqian Zhao, Rengang Li
摘要
The emergence of ChatGPT has once again sparked research in generative artificial intelligence (GAI). While people have been amazed by the generated results, they have also noticed the reasoning potential reflected in the generated textual content. However, this current ability for causal reasoning is primarily limited to the domain of language generation, such as in models like GPT-3. In visual modality, there is currently no equivalent research. Considering causal reasoning in visual content generation is significant. This is because visual information contains infinite granularity. Particularly, images can provide more intuitive and specific demonstrations for certain reasoning tasks, especially when compared to coarsegrained text. Hence, we propose a new image generation task called visual question answering with image (VQAI) and establish a dataset of the same name based on the classic Tom and Jerry animated series. Additionally, we develop a new paradigm for image generation to tackle the challenges of this task. Finally, we perform extensive experiments and analyses, including visualizations of the generated content and discussions on the potentials and limitations. The code and data are publicly available under the license of CC BY-NC-SA 4.0 for academic and non-commercial usage. The code and dataset are publicly available at: https://github.com/IEIT-AGI/MIX- Shannon/blob/main/projects/VQAI/lgd vqai.md.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Causal-Entity Reflected Egocentric Traffic Accident Video SynthesisLei-Lei Li, Jianwu Fang, Junbin Xiao, Shanmin Pang 等ICCV 2025 · 被引用 4 次
- Reasoning Diffusion for Unpaired Test Time Out-of-distribution Text-Image to Video GenerationZirui Pan, Xin Wang, Yipeng Zhang, Hong Chen 等CVPR 2026
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language ModelsMuhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Fahad KhanACL 2024 · 被引用 279 次
- Imagination Helps Visual Reasoning, But Not Yet in Latent SpaceYou Li, Chi Chen, Yanghao Li, Fanhu Zeng 等ICML 2026 · 被引用 6 次
- Transformation Driven Visual ReasoningXin Hong, Yanyan Lan, Liang Pang, Jiafeng Guo 等CVPR 2021
- Counterfactual Vision and Language LearningEhsan Abbasnejad, Damien Teney, Amin Parvaneh, Javen Shi 等CVPR 2020
- Grounded Chain-of-Thought for Multimodal Large Language ModelsQiong Wu, Xiangcong Yang, Yiyi Zhou, Chenxin Fang 等CVPR 2026 · 被引用 55 次
