ISIA Food-500: A Dataset for Large-Scale Food Recognition via Stacked Global-Local Attention Network
Weiqing Min, Linhu Liu, Zhiling Wang, Zhengdong Luo, Xiaoming Wei, Xiaolin Wei, Shuqiang Jiang
Abstract
Food recognition has received more and more attention in the multimedia community for its various real-world applications, such as diet management and self-service restaurants. A large-scale ontology of food images is urgently needed for developing advanced large-scale food recognition algorithms, as well as for providing the benchmark dataset for such algorithms. To encourage further progress in food recognition, we introduce the dataset ISIA Food-500 with 500 categories from the list in the Wikipedia and 399,726 images, a more comprehensive food dataset that surpasses existing popular benchmark datasets by category coverage and data volume. Furthermore, we propose a stacked global-local attention network, which consists of two sub-networks for food recognition. One sub-network first utilizes hybrid spatial-channel attention to extract more discriminative features, and then aggregates these multi-scale discriminative features from multiple layers into global-level representation (e.g., texture and shape information about food). The other one generates attentional regions (e.g., ingredient relevant regions) from different regions via cascaded spatial transformers, and further aggregates these multi-scale regional features from different layers into local-level representation. These two types of features are finally fused as comprehensive representation for food recognition. Extensive experiments on ISIA Food-500 and other two popular benchmark datasets demonstrate the effectiveness of our proposed method, and thus can be considered as one strong baseline. The dataset, code and models can be found at http://123.57.42.89/FoodComputing-Dataset/ISIA-Food500.html.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3099e25-ff20-440a-a138-5237b7d0408dCited by top-tier papers10
- A Large-Scale Benchmark for Food Image SegmentationXiongwei Wu, Xin Fu, Ying Liu, Ee-Peng Lim et al.ACM MM 2021 · 96 citations
- Delving into Multimodal Prompting for Fine-Grained Visual ClassificationXin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du et al.AAAI 2024 · 71 citations
- Learning Program Representations for Food Images and Cooking RecipesDim P. Papadopoulos, Enrique Mora, Nadiia Chepurko, Kuan Wei Huang et al.CVPR 2022 · 38 citations
- Learn More for Food Recognition via Progressive Self-DistillationYaohui Zhu, Linhu Liu, Jiang TianAAAI 2023 · 10 citations
- Feature-Suppressed Contrast for Self-Supervised Food Pre-trainingXinda Liu, Yaohui Zhu, Linhu Liu, Jiang Tian et al.ACM MM 2023 · 9 citations
Related papers
- Hybrid Fusion with Intra- and Cross-Modality Attention for Image-Recipe RetrievalJiao Li, Xing Xu, Wei Yu, Fumin Shen et al.SIGIR 2021 · 21 citations
- FoodLogoDet-1500: A Dataset for Large-Scale Food Logo Detection via Multi-Scale Feature Decoupling NetworkQiang Hou, Weiqing Min, Jing Wang, Sujuan Hou et al.ACM MM 2021 · 28 citations
- Visual News: Benchmark and Challenges in News Image CaptioningFuxiao Liu, Yinghan Wang, Tianlu Wang, Vicente OrdonezEMNLP 2021 · 67 citations
- Ingredients-Guided and Nutrients-Prompted Network for Food Nutrition EstimationDonglin Zhang, Boyuan Ma, Xiaojun Wu, Josef KittlerACM MM 2025 · 1 citation
- Terrace-based Food Counting and SegmentationHuu-Thanh Nguyen, Chong-Wah NgoAAAI 2021 · 9 citations
