Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim
摘要
Existing approaches for image-to-recipe retrieval have the implicit assumption that a food image can fully capture the details textually documented in its recipe. However, a food image only reflects the visual outcome of a cooked dish and not the underlying cooking process. Consequently, learning cross-modal representations to bridge the modality gap between images and recipes tends to ignore subtle, recipe-specific details that are not visually apparent but are crucial for recipe retrieval. Specifically, the representations are biased to capture the dominant visual elements, resulting in difficulty in ranking similar recipes with subtle differences in use of ingredients and cooking methods. The bias in representation learning is expected to be more severe when the training data is mixed of images and recipes sourced from different cuisines. This paper proposes a novel causal approach that predicts the culinary elements potentially overlooked in images, while explicitly injecting these elements into cross-modal representation learning to mitigate biases. Experiments are conducted on the standard monolingual Recipe1M dataset and a newly curated multilingual multicultural cuisine dataset. The results indicate that the proposed causal representation learning is capable of uncovering subtle ingredients and cooking actions and achieves impressive retrieval performance on both monolingual and multilingual multicultural datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Asymmetric Loss For Multi-Label ClassificationTal Ridnik, Emanuel Ben Baruch, Nadav Zamir, Asaf Noy 等ICCV 2021 · 被引用 778 次
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang 等SIGIR 2021 · 被引用 198 次
- Causality Inspired Representation Learning for Domain GeneralizationFangrui Lv, Jian Liang, Shuang Li, Bin Zang 等CVPR 2022 · 被引用 190 次
相关 Paper
- Cross-modal Retrieval and Synthesis (X-MRS): Closing the Modality Gap in Shared Subspace LearningRicardo Guerrero, Hai Xuan Pham, Vladimir PavlovicACM MM 2021 · 被引用 35 次
- Hybrid Fusion with Intra- and Cross-Modality Attention for Image-Recipe RetrievalJiao Li, Xing Xu, Wei Yu, Fumin Shen 等SIGIR 2021 · 被引用 21 次
- CHEF: Cross-modal Hierarchical Embeddings for Food Domain RetrievalHai Xuan Pham, Ricardo Guerrero, Vladimir Pavlovic, Jiatong LiAAAI 2021 · 被引用 22 次
- MCEN: Bridging Cross-Modal Gap between Cooking Recipes and Dish Images with Latent Variable ModelHan Fu, Rui Wu, Chenghao Liu, Jianling SunCVPR 2020
- Multi-modal Cooking Workflow Construction for Food RecipesLiangming Pan, Jingjing Chen, Jianlong Wu, Shaoteng Liu 等ACM MM 2020 · 被引用 20 次
