Multi-modal Cooking Workflow Construction for Food Recipes
Liangming Pan, Jingjing Chen, Jianlong Wu, Shaoteng Liu, Chong-Wah Ngo, Min-Yen Kan, Yu-Gang Jiang, Tat-Seng Chua
Abstract
Understanding food recipe requires anticipating the implicit causal effects of cooking actions, such that the recipe can be converted into a graph describing the temporal workflow of the recipe. This is a non-trivial task that involves common-sense reasoning. However, existing efforts rely on hand-crafted features to extract the workflow graph from recipes due to the lack of large-scale labeled datasets. Moreover, they fail to utilize the cooking images, which constitute an important part of food recipes. In this paper, we build MM-ReS, the first large-scale dataset for cooking workflow construction, consisting of 9,850 recipes with human-labeled workflow graphs. Cooking steps are multi-modal, featuring both text instructions and cooking images. We then propose a neural encoder-decoder model that utilizes both visual and textual information to construct the cooking workflow, which achieved over 20% performance gain over existing hand-crafted baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7a2b421-e34f-4163-b1cd-d332e0206860Cited by top-tier papers7
- State-aware Video Procedural CaptioningTaichi Nishimura, Atsushi Hashimoto, Yoshitaka Ushiku, Hirotaka Kameko et al.ACM MM 2021 · 15 citations
- NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video UnderstandingRunning Zhao, Zhihan Jiang, Xinchen Zhang, Chirui Chang et al.UIST 2025 · 6 citations
- Counterfactual Recipe Generation: Exploring Compositional Generalization in a Realistic ScenarioXiao Liu, Yansong Feng, Jizhi Tang, Chengang Hu et al.EMNLP 2022 · 6 citations
- Chain-of-Cooking: Cooking Process Visualization via Bidirectional Chain-of-Thought GuidanceMengling Xu, Ming Tao, Bing-Kun BaoACM MM 2025 · 1 citation
- CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image GenerationRuoxuan Zhang, Bin Wen, Hongxia Xie, Yi Yao et al.ACM MM 2025
Builds on2
Related papers
- Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe RetrievalQing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng LimACM MM 2025
- Aligning Actions Across Recipe GraphsLucia Donatelli, Theresa Schmidt, Debanjali Biswas, Arne Köhn et al.EMNLP 2021
- CHEF: Cross-modal Hierarchical Embeddings for Food Domain RetrievalHai Xuan Pham, Ricardo Guerrero, Vladimir Pavlovic, Jiatong LiAAAI 2021 · 22 citations
- CookGAN: Causality Based Text-to-Image SynthesisBin Zhu, Chong-Wah NgoCVPR 2020
- Learning Program Representations for Food Images and Cooking RecipesDim P. Papadopoulos, Enrique Mora, Nadiia Chepurko, Kuan Wei Huang et al.CVPR 2022 · 38 citations
