Adversarial Text to Continuous Image Generation
Kilichbek Haydarov, Aashiq Muhamed, Xiaoqian Shen, Jovana Lazarevic, Ivan Skorokhodov, Chamuditha Jayanga Galappaththige, Mohamed Elhoseiny
Abstract
and Technology e setting of the it is a calm and eful painting this bird is black in color with a black beak and black eye rings this bird has wings that are black and white and has a small bill hair and beard is trast against the ckground this bird is white with grey and has a long pointy beak this particular bird has a belly that is gray and has black wings CUB f riding skis on a wy surface the background is so dark and the scene is kinda gloomy but art is still good i like the setting of the snow and it is a calm and peaceful painting this bird is black in color with a black beak and black eye rings this bird has wings that are black and white and has a small bill of boys playing cer on a field the green leaves on the trees look very luscious his white hair and beard is a stark contrast against the background this bird is white with grey and has a long pointy beak this particular bird has a belly that is gray and has black wings ArtEmis CUB Figure 1. Text Conditioned Extrapolation outside of Image Boundaries: The red rectangles indicate the resolution boundaries that our HyperCGAN model was trained. By design, our model can synthesize meaningful pixels at surrounding (x, y) coordinates beyond these boundaries without any explicit training. For example, it can meaningfully extend bird images with more natural details like the tail, background, and the branch of the tree.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation DetectionJinjie Shen, Yaxiong Wang, Yujiao Wu, Lechao Cheng et al.ICML 2026
- OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RLJinjie Shen, Jing Wu, Yaxiong Wang, Lechao Cheng et al.ICML 2026
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Unite and Conquer: Plug & Play Multi-Modal Synthesis Using Diffusion ModelsNithin Gopalakrishnan Nair, Wele Gedara Chaminda Bandara, Vishal M. PatelCVPR 2023
- GLIGEN: Open-Set Grounded Text-to-Image GenerationYuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu et al.CVPR 2023
- ArtAdapter: Text-to-Image Style Transfer using Multi-Level Style Encoder and Explicit AdaptationDar-Yen Chen, Hamish Tennent, Ching-Wen HsuCVPR 2024
- Exploring Sparse MoE in GANs for Text-conditioned Image SynthesisJiapeng Zhu, Ceyuan Yang, Kecheng Zheng, Yinghao Xu et al.CVPR 2025
- Democratizing Fine-grained Visual Recognition with Large Language ModelsMingxuan Liu, Subhankar Roy, Wenjing Li, Zhun Zhong et al.ICLR 2024 · 27 citations
