ISIA Food-500: A Dataset for Large-Scale Food Recognition via Stacked Global-Local Attention Network
Weiqing Min, Linhu Liu, Zhiling Wang, Zhengdong Luo, Xiaoming Wei, Xiaolin Wei, Shuqiang Jiang
摘要
Food recognition has received more and more attention in the multimedia community for its various real-world applications, such as diet management and self-service restaurants. A large-scale ontology of food images is urgently needed for developing advanced large-scale food recognition algorithms, as well as for providing the benchmark dataset for such algorithms. To encourage further progress in food recognition, we introduce the dataset ISIA Food-500 with 500 categories from the list in the Wikipedia and 399,726 images, a more comprehensive food dataset that surpasses existing popular benchmark datasets by category coverage and data volume. Furthermore, we propose a stacked global-local attention network, which consists of two sub-networks for food recognition. One sub-network first utilizes hybrid spatial-channel attention to extract more discriminative features, and then aggregates these multi-scale discriminative features from multiple layers into global-level representation (e.g., texture and shape information about food). The other one generates attentional regions (e.g., ingredient relevant regions) from different regions via cascaded spatial transformers, and further aggregates these multi-scale regional features from different layers into local-level representation. These two types of features are finally fused as comprehensive representation for food recognition. Extensive experiments on ISIA Food-500 and other two popular benchmark datasets demonstrate the effectiveness of our proposed method, and thus can be considered as one strong baseline. The dataset, code and models can be found at http://123.57.42.89/FoodComputing-Dataset/ISIA-Food500.html.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- A Large-Scale Benchmark for Food Image SegmentationXiongwei Wu, Xin Fu, Ying Liu, Ee-Peng Lim 等ACM MM 2021 · 被引用 96 次
- Delving into Multimodal Prompting for Fine-Grained Visual ClassificationXin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du 等AAAI 2024 · 被引用 71 次
- Learning Program Representations for Food Images and Cooking RecipesDim P. Papadopoulos, Enrique Mora, Nadiia Chepurko, Kuan Wei Huang 等CVPR 2022 · 被引用 38 次
- Learn More for Food Recognition via Progressive Self-DistillationYaohui Zhu, Linhu Liu, Jiang TianAAAI 2023 · 被引用 10 次
- Feature-Suppressed Contrast for Self-Supervised Food Pre-trainingXinda Liu, Yaohui Zhu, Linhu Liu, Jiang Tian 等ACM MM 2023 · 被引用 9 次
相关 Paper
- Hybrid Fusion with Intra- and Cross-Modality Attention for Image-Recipe RetrievalJiao Li, Xing Xu, Wei Yu, Fumin Shen 等SIGIR 2021 · 被引用 21 次
- FoodLogoDet-1500: A Dataset for Large-Scale Food Logo Detection via Multi-Scale Feature Decoupling NetworkQiang Hou, Weiqing Min, Jing Wang, Sujuan Hou 等ACM MM 2021 · 被引用 28 次
- Visual News: Benchmark and Challenges in News Image CaptioningFuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente OrdonezEMNLP 2021 · 被引用 67 次
- Ingredients-Guided and Nutrients-Prompted Network for Food Nutrition EstimationDonglin Zhang, Boyuan Ma, Xiaojun Wu, Josef KittlerACM MM 2025 · 被引用 1 次
- Terrace-based Food Counting and SegmentationHuu-Thanh Nguyen, Chong-Wah NgoAAAI 2021 · 被引用 9 次
