ChefGAN: Food Image Generation from Recipes
Siyuan Pan, Ling Dai, Xuhong Hou, Huating Li, Bin Sheng
Abstract
Although significant progress has been made in generating images from the text by using generative adversarial networks (GANs), it is still challenging to deal with long text, which contains complex semantic information like recipes. This paper focuses on generating images with high visual realism and semantic consistency from the complex text of recipes. To achieve this, we propose a GANs based method termed ChefGAN. The critical concept of ChefGAN is that a joint image-recipe embedding model is used before the generation task to provide high-quality representations of recipes, and it acts as an extra regularization during the generation to improve semantic consistency. Two modules are designed for this image text embedding module (ITEM) and a cascaded image generation module (CIGM). The generation process is carried out in 3 steps: (1) Two encoders in ITEM are trained simultaneously to generate similar representations for each image-recipe pair. (2) CIGM generates images according to the representations from ITEM's text encoder. (3) The generated image is fed into ITEM's image encoder to calculate the similarity with the given recipe. This process can provide additional regularization effect other than the impact of a discriminator. To facilitate convergence, we applied a two-stage training strategy, which generates an image with low resolution and then one with high resolution in the CIGM module. Compared with other representative state-of-the-art methods, ChefGAN demonstrates better performance both in visual realism and semantic consistency.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d2450203-4253-4abf-bf94-c674853acf89Cited by top-tier papers5
- Cross-Modal Recipe Embeddings by Disentangling Recipe Contents and Dish StylesYu Sugiyama, Keiji YanaiACM MM 2021 · 15 citations
- Navigating Weight Prediction with Diet DiaryYinxuan Gui, Bin Zhu, Jingjing Chen, Chong Wah Ngo et al.ACM MM 2024 · 6 citations
- Chain-of-Cooking: Cooking Process Visualization via Bidirectional Chain-of-Thought GuidanceMengling Xu, Ming Tao, Bing-Kun BaoACM MM 2025 · 1 citation
- CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image GenerationRuoxuan Zhang, Bin Wen, Hongxia Xie, Yi Yao et al.ACM MM 2025
- Revamping Cross-Modal Recipe Retrieval With Hierarchical Transformers and Self-Supervised LearningAmaia Salvador, Erhan Gundogdu, Loris Bazzani, Michael DonoserCVPR 2021
Related papers
- CookGAN: Causality Based Text-to-Image SynthesisBin Zhu, Chong-Wah NgoCVPR 2020
- Cycle-Consistent Inverse GAN for Text-to-Image SynthesisHao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan MiaoACM MM 2021 · 47 citations
- Learning Program Representations for Food Images and Cooking RecipesDim P. Papadopoulos, Enrique Mora, Nadiia Chepurko, Kuan Wei Huang et al.CVPR 2022 · 38 citations
- DF-GAN: A Simple and Effective Baseline for Text-to-Image SynthesisMing Tao, Hao Tang, Fei Wu, Xiaoyuan Jing et al.CVPR 2022 · 296 citations
- IR-GAN: Image Manipulation with Linguistic Instruction by Increment ReasoningZhenhuan Liu, Jincan Deng, Liang Li, Shaofei Cai et al.ACM MM 2020 · 17 citations
