MCEN: Bridging Cross-Modal Gap between Cooking Recipes and Dish Images with Latent Variable Model
Han Fu, Rui Wu, Chenghao Liu, Jianling Sun
摘要
Nowadays, driven by the increasing concern on diet and health, food computing has attracted enormous attention from both industry and research community. One of the most popular research topics in this domain is Food Retrieval, due to its profound influence on health-oriented applications. In this paper, we focus on the task of cross-modal retrieval between food images and cooking recipes. We present Modality-Consistent Embedding Network (MCEN) that learns modality-invariant representations by projecting images and texts to the same embedding space. To capture the latent alignments between modalities, we incorporate stochastic latent variables to explicitly exploit the interactions between textual and visual features. Importantly, our method learns the cross-modal alignments during training but computes embeddings of different modalities independently at inference time for the sake of efficiency. Extensive experimental results clearly demonstrate that the proposed MCEN outperforms all existing approaches on the benchmark Recipe1M dataset and requires less computational cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Learning Program Representations for Food Images and Cooking RecipesDim P. Papadopoulos, Enrique Mora, Nadiia Chepurko, Kuan Wei Huang 等CVPR 2022 · 被引用 38 次
- Cross-modal Retrieval and Synthesis (X-MRS): Closing the Modality Gap in Shared Subspace LearningRicardo Guerrero, Hai Xuan Pham, Vladimir PavlovicACM MM 2021 · 被引用 35 次
- Cross-Modal Recipe Embeddings by Disentangling Recipe Contents and Dish StylesYu Sugiyama, Keiji YanaiACM MM 2021 · 被引用 15 次
- Paired Cross-Modal Data Augmentation for Fine-Grained Image-to-Text RetrievalHao Wang, Guosheng Lin, Steven C. H. Hoi, Chunyan MiaoACM MM 2022 · 被引用 12 次
- Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe RetrievalQing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng LimACM MM 2025
它引用的顶会 Paper2
相关 Paper
- CHEF: Cross-modal Hierarchical Embeddings for Food Domain RetrievalHai Xuan Pham, Ricardo Guerrero, Vladimir Pavlovic, Jiatong LiAAAI 2021 · 被引用 22 次
- Hybrid Fusion with Intra- and Cross-Modality Attention for Image-Recipe RetrievalJiao Li, Xing Xu, Wei Yu, Fumin Shen 等SIGIR 2021 · 被引用 21 次
- Revamping Cross-Modal Recipe Retrieval With Hierarchical Transformers and Self-Supervised LearningAmaia Salvador, Erhan Gundogdu, Loris Bazzani, Michael DonoserCVPR 2021
- Improving Cross-Modal Recipe Retrieval with Component-Aware Prompted CLIP EmbeddingXu Huang, Jin Liu, Zhizhong Zhang, Yuan XieACM MM 2023 · 被引用 11 次
- MCCN: Multimodal Coordinated Clustering Network for Large-Scale Cross-modal RetrievalZhixiong Zeng, Ying Sun, Wenji MaoACM MM 2021 · 被引用 20 次
