NutriBench: A Dataset for Evaluating Large Language Models in Nutrition Estimation from Meal Descriptions
Mehak Preet Dhaliwal, Andong Hua, Laya Pullela, Ryan Burke, Yao Qin
Abstract
Accurate nutrition estimation helps people make informed dietary choices and is essential in the prevention of serious health complications. We present NU-TRIBENCH, the first publicly available natural language meal description nutrition benchmark. NUTRIBENCH consists of 11,857 meal descriptions generated from real-world global dietary intake data. The data is human-verified and annotated with macro-nutrient labels, including carbohydrates, proteins, fats, and calories. We conduct an extensive evaluation of NUTRIBENCH on the task of carbohydrate estimation, testing twelve leading Large Language Models (LLMs), including GPT-4o, Llama3.1, Qwen2, Gemma2, and OpenBioLLM models, using standard, Chain-of-Thought and Retrieval-Augmented Generation strategies. Additionally, we present a study involving professional nutritionists, finding that LLMs can provide comparable but significantly faster estimates. Finally, we perform a real-world risk assessment by simulating the effect of carbohydrate predictions on the blood glucose levels of individuals with diabetes. Our work highlights the opportunities and challenges of using LLMs for nutrition estimation, demonstrating their potential to aid professionals and laypersons and improve health outcomes. Our benchmark is publicly available at: https://mehak126.github.io/nutribench.html * Equal contribution, alphabetically ordered.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83b11af7-a3a2-4ec3-8e15-d602e439e821Builds on8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang et al.ICLR 2020 · 674 citations
Related papers
- NGQA: A Nutritional Graph Question Answering Benchmark for Personalized Health-aware Nutritional ReasoningZheyuan Zhang, Yiyang Li, Nhi Ha Lan Le, Zehong Wang et al.ACL 2025 · 18 citations
- VisNumBench: Evaluating Number Sense of Multimodal Large Language ModelsTengjin Weng, Jingyi Wang, Wenhao Jiang, Zhong MingICCV 2025 · 1 citation
- Integrating Expertise in LLMs: Crafting a Customized Nutrition Assistant with Refined Template InstructionsAnnalisa Szymanski, Brianna L. Wimer, Oghenemaro Anuyah, Heather A. Eicher-Miller et al.CHI 2024 · 27 citations
- MedAraBench: Large-scale Arabic Medical Question Answering Dataset and BenchmarkMouath Abu Daoud, Leen Kharouf, Omar El Hajj, Dana El Samad et al.ICLR 2026 · 4 citations
- DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary DomainSong Jin, Juntian Zhang, Xun Zhang, Zeying Tian et al.ACL 2026 · 1 citation
