CHEF: Cross-modal Hierarchical Embeddings for Food Domain Retrieval
Hai Xuan Pham, Ricardo Guerrero, Vladimir Pavlovic, Jiatong Li
摘要
Despite the abundance of multi-modal data, such as image-text pairs, there has been little effort in understanding the individual entities and their different roles in the construction of these data instances. In this work, we endeavour to discover the entities and their corresponding importance in cooking recipes automatically as a visual-linguistic association problem. More specifically, we introduce a novel cross-modal learning framework to jointly model the latent representations of images and text in the food image-recipe association and retrieval tasks. This model allows one to discover complex functional and hierarchical relationships between images and text, and among textual parts of a recipe including title, ingredients and cooking instructions. Our experiments show that by making use of efficient tree-structured Long Short-Term Memory as the text encoder in our computational cross-modal retrieval framework, we are not only able to identify the main ingredients and cooking actions in the recipe descriptions without explicit supervision, but we can also learn more meaningful feature representations of food recipes, appropriate for challenging cross-modal retrieval and recipe adaption tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Learning Program Representations for Food Images and Cooking RecipesDim P. Papadopoulos, Enrique Mora, Nadiia Chepurko, Kuan Wei Huang 等CVPR 2022 · 被引用 38 次
- Cross-modal Retrieval and Synthesis (X-MRS): Closing the Modality Gap in Shared Subspace LearningRicardo Guerrero, Hai Xuan Pham, Vladimir PavlovicACM MM 2021 · 被引用 35 次
相关 Paper
- Revamping Cross-Modal Recipe Retrieval With Hierarchical Transformers and Self-Supervised LearningAmaia Salvador, Erhan Gundogdu, Loris Bazzani, Michael DonoserCVPR 2021
- Hybrid Fusion with Intra- and Cross-Modality Attention for Image-Recipe RetrievalJiao Li, Xing Xu, Wei Yu, Fumin Shen 等SIGIR 2021 · 被引用 21 次
- MCEN: Bridging Cross-Modal Gap between Cooking Recipes and Dish Images with Latent Variable ModelHan Fu, Rui Wu, Chenghao Liu, Jianling SunCVPR 2020
- Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe RetrievalQing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng LimACM MM 2025
- Improving Cross-Modal Recipe Retrieval with Component-Aware Prompted CLIP EmbeddingXu Huang, Jin Liu, Zhizhong Zhang, Yuan XieACM MM 2023 · 被引用 11 次
