Multi-modal Cooking Workflow Construction for Food Recipes
Liangming Pan, Jingjing Chen, Jianlong Wu, Shaoteng Liu, Chong-Wah Ngo, Min-Yen Kan, Yu-Gang Jiang, Tat-Seng Chua
摘要
Understanding food recipe requires anticipating the implicit causal effects of cooking actions, such that the recipe can be converted into a graph describing the temporal workflow of the recipe. This is a non-trivial task that involves common-sense reasoning. However, existing efforts rely on hand-crafted features to extract the workflow graph from recipes due to the lack of large-scale labeled datasets. Moreover, they fail to utilize the cooking images, which constitute an important part of food recipes. In this paper, we build MM-ReS, the first large-scale dataset for cooking workflow construction, consisting of 9,850 recipes with human-labeled workflow graphs. Cooking steps are multi-modal, featuring both text instructions and cooking images. We then propose a neural encoder-decoder model that utilizes both visual and textual information to construct the cooking workflow, which achieved over 20% performance gain over existing hand-crafted baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- State-aware Video Procedural CaptioningTaichi Nishimura, Atsushi Hashimoto, Yoshitaka Ushiku, Hirotaka Kameko 等ACM MM 2021 · 被引用 15 次
- NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video UnderstandingRunning Zhao, Zhihan Jiang, Xinchen Zhang, Chirui Chang 等UIST 2025 · 被引用 6 次
- Counterfactual Recipe Generation: Exploring Compositional Generalization in a Realistic ScenarioXiao Liu, Yansong Feng, Jizhi Tang, Chengang Hu 等EMNLP 2022 · 被引用 6 次
- Chain-of-Cooking: Cooking Process Visualization via Bidirectional Chain-of-Thought GuidanceMengling Xu, Ming Tao, Bing-Kun BaoACM MM 2025 · 被引用 1 次
- CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image GenerationRuoxuan Zhang, Bin Wen, Hongxia Xie, Yi Yao 等ACM MM 2025
它引用的顶会 Paper2
相关 Paper
- Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe RetrievalQing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng LimACM MM 2025
- Aligning Actions Across Recipe GraphsLucia Donatelli, Theresa Schmidt, Debanjali Biswas, Arne Köhn 等EMNLP 2021
- CHEF: Cross-modal Hierarchical Embeddings for Food Domain RetrievalHai Xuan Pham, Ricardo Guerrero, Vladimir Pavlovic, Jiatong LiAAAI 2021 · 被引用 22 次
- CookGAN: Causality Based Text-to-Image SynthesisBin Zhu, Chong-Wah NgoCVPR 2020
- Learning Program Representations for Food Images and Cooking RecipesDim P. Papadopoulos, Enrique Mora, Nadiia Chepurko, Kuan Wei Huang 等CVPR 2022 · 被引用 38 次
