Revamping Cross-Modal Recipe Retrieval With Hierarchical Transformers and Self-Supervised Learning
Amaia Salvador, Erhan Gundogdu, Loris Bazzani, Michael Donoser
摘要
Cross-modal recipe retrieval has recently gained substantial attention due to the importance of food in people's lives, as well as the availability of vast amounts of digital cooking recipes and food images to train machine learning models. In this work, we revisit existing approaches for cross-modal recipe retrieval and propose a simplified end-to-end model based on well established and high performing encoders for text and images. We introduce a hierarchical recipe Transformer which attentively encodes individual recipe components (titles, ingredients and instructions). Further, we propose a self-supervised loss function computed on top of pairs of individual recipe components, which is able to leverage semantic relationships within recipes, and enables training using both image-recipe and recipe-only samples. We conduct a thorough analysis and ablation studies to validate our design choices. As a result, our proposed method achieves state-of-the-art performance in the cross-modal recipe retrieval task on the Recipe1M dataset. We make code and models publicly available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- All in One: Exploring Unified Vision-Language Tracking with Multi-Modal AlignmentChunhui Zhang, Xin Sun, Yiqian Yang, Li Liu 等ACM MM 2023 · 被引用 41 次
- Learning Program Representations for Food Images and Cooking RecipesDim P. Papadopoulos, Enrique Mora, Nadiia Chepurko, Kuan Wei Huang 等CVPR 2022 · 被引用 38 次
- Cross-modal Retrieval and Synthesis (X-MRS): Closing the Modality Gap in Shared Subspace LearningRicardo Guerrero, Hai Xuan Pham, Vladimir PavlovicACM MM 2021 · 被引用 35 次
- Review-Enhanced Hierarchical Contrastive Learning for RecommendationKe Wang, Yanmin Zhu, Tianzi Zang, Chunyang Wang 等AAAI 2024 · 被引用 17 次
- Paired Cross-Modal Data Augmentation for Fine-Grained Image-to-Text RetrievalHao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan MiaoACM MM 2022 · 被引用 12 次
它引用的顶会 Paper6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- ISIA Food-500: A Dataset for Large-Scale Food Recognition via Stacked Global-Local Attention NetworkWeiqing Min, Linhu Liu, Zhiling Wang, Zhengdong Luo 等ACM MM 2020 · 被引用 154 次
- ACMM: Aligned Cross-Modal Memory for Few-Shot Image and Sentence MatchingYan Huang, Liang WangICCV 2019 · 被引用 68 次
- ChefGAN: Food Image Generation from RecipesSiyuan Pan, Ling Dai, Xuhong Hou, Huating Li 等ACM MM 2020 · 被引用 32 次
相关 Paper
- CHEF: Cross-modal Hierarchical Embeddings for Food Domain RetrievalHai Xuan Pham, Ricardo Guerrero, Vladimir Pavlovic, Jiatong LiAAAI 2021 · 被引用 22 次
- Hybrid Fusion with Intra- and Cross-Modality Attention for Image-Recipe RetrievalJiao Li, Xing Xu, Wei Yu, Fumin Shen 等SIGIR 2021 · 被引用 21 次
- Improving Cross-Modal Recipe Retrieval with Component-Aware Prompted CLIP EmbeddingXu Huang, Jin Liu, Zhizhong Zhang, Yuan XieACM MM 2023 · 被引用 11 次
- MCEN: Bridging Cross-Modal Gap between Cooking Recipes and Dish Images with Latent Variable ModelHan Fu, Rui Wu, Chenghao Liu, Jianling SunCVPR 2020
- Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe RetrievalQing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng LimACM MM 2025
