Hybrid Fusion with Intra- and Cross-Modality Attention for Image-Recipe Retrieval
Jiao Li, Xing Xu, Wei Yu, Fumin Shen, Zuo Cao, Kai Zuo, Heng Tao Shen
摘要
Image-recipe retrieval, which aims at retrieving the relevant recipe from a food image and vice versa, is now attracting widespread attention, since sharing food-related images and recipes on the Internet has become a popular trend. Existing methods have formulated this problem as a typical cross-modal retrieval task by learning the image-recipe similarity. Though these methods have made inspiring achievements for image-recipe retrieval, they may still be less effective to jointly incorporate the three crucial points: (1) the association between ingredients and instructions, (2) fine-grained image information, and (3) the latent alignment between recipes and images. To this end, we propose a novel framework namedHybrid Fusion with Intra- and Cross-Modality Attention (HF-ICMA) to learn accurate image-recipe similarity. Our HF-ICMA model adopts an intra-recipe fusion module to focus on the interaction between ingredients and instructions within a recipe, and further enriches the expressions of the two separate embeddings. Meanwhile, an image-recipe fusion module is devised to explore the potential relationship between fine-grained image regions and ingredients from the recipe, which jointly forms the final image-recipe similarity from both the local and global aspects. Extensive experiments on the large-scale benchmark dataset Recipe1M show that our model significantly outperforms the state-of-the-art approaches on various image-recipe retrieval scenarios.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- Dual Semantic Knowledge Composed Multimodal Dialog SystemsXiaolin Chen, Xuemeng Song, Yinwei Wei, Liqiang Nie 等SIGIR 2023 · 被引用 8 次
- Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe RetrievalQing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng LimACM MM 2025
- Multimodal Lego: Model Merging and Fine-Tuning Across Topologies and Modalities in BiomedicineKonstantin Hemker, Nikola Simidjievski, Mateja JamnikICLR 2025
相关 Paper
- Improving Cross-Modal Recipe Retrieval with Component-Aware Prompted CLIP EmbeddingXu Huang, Jin Liu, Zhizhong Zhang, Yuan XieACM MM 2023 · 被引用 11 次
- Revamping Cross-Modal Recipe Retrieval With Hierarchical Transformers and Self-Supervised LearningAmaia Salvador, Erhan Gundogdu, Loris Bazzani, Michael DonoserCVPR 2021
- CHEF: Cross-modal Hierarchical Embeddings for Food Domain RetrievalHai Xuan Pham, Ricardo Guerrero, Vladimir Pavlovic, Jiatong LiAAAI 2021 · 被引用 22 次
- MCEN: Bridging Cross-Modal Gap between Cooking Recipes and Dish Images with Latent Variable ModelHan Fu, Rui Wu, Chenghao Liu, Jianling SunCVPR 2020
- ISIA Food-500: A Dataset for Large-Scale Food Recognition via Stacked Global-Local Attention NetworkWeiqing Min, Linhu Liu, Zhiling Wang, Zhengdong Luo 等ACM MM 2020 · 被引用 154 次
