A Recipe for Creating Multimodal Aligned Datasets for Sequential Tasks
Angela S. Lin, Sudha Rao, Asli Celikyilmaz, Elnaz Nouri, Chris Brockett, Debadeepta Dey, Bill Dolan
摘要
Many high-level procedural tasks can be decomposed into sequences of instructions that vary in their order and choice of tools. In the cooking domain, the web offers many partially-overlapping text and video recipes (i.e. procedures) that describe how to make the same dish (i.e. high-level task). Aligning instructions for the same dish across different sources can yield descriptive visual explanations that are far richer semantically than conventional textual instructions, providing commonsense insight into how real-world procedures are structured. Learning to align these different instruction sets is challenging because: a) different recipes vary in their order of instructions and use of ingredients; and b) video instructions can be noisy and tend to contain far more information than text instructions. To address these challenges, we first use an unsupervised alignment algorithm that learns pairwise alignments between instructions of different recipes for the same dish. We then use a graph algorithm to derive a joint alignment between multiple text and multiple video recipes for the same dish. We release the MICROSOFT RESEARCH MUL-TIMODAL ALIGNED RECIPE CORPUS 1 containing ∼150K pairwise alignments between recipes across 4,262 dishes with rich commonsense information. * Work done when the author was an intern at Microsoft. 1 https://github.com/microsoft/ multimodal-aligned-recipe-corpus 7. Add 12 ounces of thawed peas and bean sprouts. 3. Add onion, garlic, peas and carrots. 4. Transfer shrimp to the hot skillet and cook them one minute per side.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- : Visualization of AI-Assisted Task Guidance in ARSonia Castelo, João Rulff, Erin McGowan, Bea Steers 等IEEE VIS 2023 · 被引用 33 次
- Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine ComprehensionHuibin Zhang, Zhengkun Zhang, Yao Zhang, Jun Wang 等ACL 2022 · 被引用 5 次
- Substance over Style: Document-Level Targeted Content TransferAllison Hegel, Sudha Rao, Asli Celikyilmaz, Bill DolanEMNLP 2020
- CaT-Bench: Benchmarking Language Model Understanding of Causal and Temporal Dependencies in PlansYash Kumar Lal, Vanya Cohen, Nathanael Chambers, Niranjan Balasubramanian 等EMNLP 2024
- Aligning Actions Across Recipe GraphsLucia Donatelli, Theresa Schmidt, Debanjali Biswas, Arne Köhn 等EMNLP 2021
相关 Paper
- Zero-Shot Anticipation for Instructional ActivitiesFadime Sener, Angela YaoICCV 2019 · 被引用 75 次
- GePSAn: Generative Procedure Step Anticipation in Cooking VideosMohamed Ashraf Abdelsalam, Samrudhdhi B. Rangrej, Isma Hadji, Nikita Dvornik 等ICCV 2023 · 被引用 10 次
- Learning Program Representations for Food Images and Cooking RecipesDim P. Papadopoulos, Enrique Mora, Nadiia Chepurko, Kuan Wei Huang 等CVPR 2022 · 被引用 38 次
- Multi-modal Cooking Workflow Construction for Food RecipesLiangming Pan, Jingjing Chen, Jianlong Wu, Shaoteng Liu 等ACM MM 2020 · 被引用 20 次
- Stitch-a-Demo: Creating Video Demonstrations from Multistep DescriptionsChi Hsuan Wu, Kumar Ashutosh, Kristen GraumanCVPR 2026 · 被引用 1 次
