OmniFood8K: Single-Image Nutrition Estimation via Hierarchical Frequency-Aligned Fusion
Dongjian Yu, Weiqing Min, Qian Jiang, Xing Lin, Xin Jin, Shuqiang Jiang
摘要
Accurate estimation of food nutrition plays a vital role in promoting healthy dietary habits and personalized diet management. Most existing food datasets primarily focus on Western cuisines and lack sufficient coverage of Chinese dishes, which restricts accurate nutritional estimation for Chinese meals. Moreover, many state-of-theart nutrition prediction methods rely on depth sensors, restricting their applicability in daily scenarios. To address these limitations, we introduce OmniFood8K, a comprehensive multimodal dataset comprising 8,036 food samples, each with detailed nutritional annotations and multiview images. In addition, to enhance models' capability in nutritional prediction, we construct NutritionSynth-115K, a large-scale synthetic dataset that introduces compositional variations while preserving precise nutritional labels. Moreover, we propose an end-to-end framework for nutritional prediction from a single RGB image. First, we predict a depth map from a single RGB image and design the Scale-Shift Residual Adapter (SSRA) to refine it for global scale consistency and local structural preservation. Second, we propose the Frequency-Aligned Fusion Module (FAFM) to hierarchically align and fuse RGB and depth features in the frequency domain. Finally, we design a Mask-based Prediction Head (MPH) to emphasize key ingredient regions via dynamic channel selection for more accurate prediction. Extensive experiments on multiple datasets demonstrate the superiority of our method over existing approaches. Project homepage: https://yudongjian.github.io/OmniFood8K-food/ * Corresponding author Mass Calories Protein Fat Carb. Total 206.0 2729.5 19.0 55.9 19.7 Garlic scape 102.9 308.7 2.0 0.1 15.8
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- TriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding NetworkZhengyi Liu, Yuan Wang, Zhengzheng Tu, Yun Xiao 等ACM MM 2021 · 被引用 175 次
相关 Paper
- Nutrition5k: Towards Automatic Nutritional Understanding of Generic FoodQuin Thames, Arjun Karpur, Wade Norris, Fangting Xia 等CVPR 2021
- Ingredients-Guided and Nutrients-Prompted Network for Food Nutrition EstimationDonglin Zhang, Boyuan Ma, Xiaojun Wu, Josef KittlerACM MM 2025 · 被引用 1 次
- Spatial-Aware Multi-Modal Information Fusion for Food Nutrition EstimationDongjian Yu, Weiqing Min, Xin Jin, Qian Jiang 等ACM MM 2025 · 被引用 1 次
- A Large-Scale Benchmark for Food Image SegmentationXiongwei Wu, Xin Fu, Ying Liu, Ee-Peng Lim 等ACM MM 2021 · 被引用 96 次
- DSDGF-Nutri: A Decoupled Self-Distillation Network with Gating Fusion For Food Nutritional AssessmentSujuan Hou, Zhihui Feng, Hao Xiong, Weiqing Min 等ACM MM 2025 · 被引用 2 次
