A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task
Mashiro Toyooka, Kiyoharu Aizawa, Yoko Yamakata
Abstract
Large Language Models (LLMs) are trained on a vast amount of procedural texts, but they do not directly observe real-world phenomena. In the context of cooking recipes, this poses a challenge, as intermediate states of ingredients are often omitted, making it difficult for models to track ingredient states and understand recipes accurately. In this paper, we apply state probing, a method for evaluating a language model's understanding of the world, to the domain of cooking. We propose a new task and dataset for evaluating how well LLMs can recognize intermediate ingredient states during cooking procedures. We first construct a new Japanese recipe dataset with clear and accurate annotations of ingredient state changes, collected from well-structured and controlled recipe texts. Using this dataset, we design three novel tasks to evaluate whether LLMs can track ingredient state transitions and identify ingredients present at intermediate steps. Our experiments with widely used LLMs, such as Llama3.1-70B and Qwen2.5-72B, show that learning ingredient state knowledge improves their understanding of cooking processes, achieving performance comparable to commercial LLMs. The dataset are publicly available at: https://huggingface.co/datasets/mashi6n/nhkrecipe-100-anno-1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b65fa7ac-2e72-45b7-8f7a-9d03752217dcCited by top-tier papers1
Ask how each one uses itBuilds on6
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Learning Program Representations for Food Images and Cooking RecipesDim P. Papadopoulos, Enrique Mora, Nadiia Chepurko, Kuan Wei Huang et al.CVPR 2022 · 38 citations
- Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional ManualsTe-Lin Wu, Alexander Spangher, Pegah Alipoormolabashi, Marjorie Freedman et al.ACL 2022 · 30 citations
- Multi-modal Cooking Workflow Construction for Food RecipesLiangming Pan, Jingjing Chen, Jianlong Wu, Shaoteng Liu et al.ACM MM 2020 · 20 citations
Related papers
- Counterfactual Recipe Generation: Exploring Compositional Generalization in a Realistic ScenarioXiao Liu, Yansong Feng, Jizhi Tang, Chengang Hu et al.EMNLP 2022 · 6 citations
- Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative AgentsHaochen Sun, Shuwen Zhang, Lujie Niu, Lei Ren et al.EMNLP 2025
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?Apratim Bhattacharyya, Bicheng Xu, Sanjay Haresh, Reza Pourreza et al.NeurIPS 2025 · 2 citations
- Predicting Implicit Arguments in Procedural Video InstructionsAnil Batra, Laura Sevilla-Lara, Marcus Rohrbach, Frank KellerACL 2025
- PAGED: A Benchmark for Procedural Graphs Extraction from DocumentsWeihong Du, Wenrui Liao, Hongru Liang, Wenqiang LeiACL 2024 · 4 citations
