Learning Program Representations for Food Images and Cooking Recipes
Dim P. Papadopoulos, Enrique Mora, Nadiia Chepurko, Kuan Wei Huang, Ferda Ofli, Antonio Torralba
Abstract
In this paper, we are interested in modeling a how-to instructional procedure, such as a cooking recipe, with a meaningful and rich high-level representation. Specifically, we propose to represent cooking recipes and food images as cooking programs. Programs provide a structured repre-sentation of the task, capturing cooking semantics and se-quential relationships of actions in the form of a graph. This allows them to be easily manipulated by users and executed by agents. To this end, we build a model that is trained to learn a joint embedding between recipes and food images via self-supervision and jointly generate a program from this embedding as a sequence. To validate our idea, we crowdsource programs for cooking recipes and show that: (a) projecting the image-recipe embeddings into programs leads to better cross-modal retrieval results; (b) generating programs from images leads to better recognition re-sults compared to predicting raw cooking instructions; and (c) we can generate food images by manipulating programs via optimizing the latent code of a GAN. Code, data, and models are available online <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> http://cookingprograms.csail.mit.edu.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd63c8a2-f638-4d7c-ae57-bf8653577323Cited by top-tier papers4
- Clover: Toward Sustainable AI with Carbon-Aware Machine Learning Inference ServiceBaolin Li, Siddharth Samsi, Vijay Gadepally, Devesh TiwariSC 2023 · 63 citations
- Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language NavigationXiang Fang, Wanlong Fang, Changshuo WangNeurIPS 2025 · 24 citations
- Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe RetrievalQing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng LimACM MM 2025
- A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing TaskMashiro Toyooka, Kiyoharu Aizawa, Yoko YamakataACM MM 2025
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
- ISIA Food-500: A Dataset for Large-Scale Food Recognition via Stacked Global-Local Attention NetworkWeiqing Min, Linhu Liu, Zhiling Wang, Zhengdong Luo et al.ACM MM 2020 · 154 citations
Related papers
- RECIPTOR: An Effective Pretrained Model for Recipe Representation LearningDiya Li, Mohammed J. ZakiKDD 2020 · 28 citations
- ChefGAN: Food Image Generation from RecipesSiyuan Pan, Ling Dai, Xuhong Hou, Huating Li et al.ACM MM 2020 · 32 citations
- Cross-Modal Recipe Embeddings by Disentangling Recipe Contents and Dish StylesYu Sugiyama, Keiji YanaiACM MM 2021 · 15 citations
- Multi-modal Cooking Workflow Construction for Food RecipesLiangming Pan, Jingjing Chen, Jianlong Wu, Shaoteng Liu et al.ACM MM 2020 · 20 citations
- MCEN: Bridging Cross-Modal Gap between Cooking Recipes and Dish Images with Latent Variable ModelHan Fu, Rui Wu, Chenghao Liu, Jianling SunCVPR 2020
