NutriBench: A Dataset for Evaluating Large Language Models in Nutrition Estimation from Meal Descriptions
Mehak Preet Dhaliwal, Andong Hua, Laya Pullela, Ryan Burke, Yao Qin
摘要
Accurate nutrition estimation helps people make informed dietary choices and is essential in the prevention of serious health complications. We present NU-TRIBENCH, the first publicly available natural language meal description nutrition benchmark. NUTRIBENCH consists of 11,857 meal descriptions generated from real-world global dietary intake data. The data is human-verified and annotated with macro-nutrient labels, including carbohydrates, proteins, fats, and calories. We conduct an extensive evaluation of NUTRIBENCH on the task of carbohydrate estimation, testing twelve leading Large Language Models (LLMs), including GPT-4o, Llama3.1, Qwen2, Gemma2, and OpenBioLLM models, using standard, Chain-of-Thought and Retrieval-Augmented Generation strategies. Additionally, we present a study involving professional nutritionists, finding that LLMs can provide comparable but significantly faster estimates. Finally, we perform a real-world risk assessment by simulating the effect of carbohydrate predictions on the blood glucose levels of individuals with diabetes. Our work highlights the opportunities and challenges of using LLMs for nutrition estimation, demonstrating their potential to aid professionals and laypersons and improve health outcomes. Our benchmark is publicly available at: https://mehak126.github.io/nutribench.html * Equal contribution, alphabetically ordered.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
相关 Paper
- NGQA: A Nutritional Graph Question Answering Benchmark for Personalized Health-aware Nutritional ReasoningZheyuan Zhang, Yiyang Li, Nhi Ha Lan Le, Zehong Wang 等ACL 2025 · 被引用 18 次
- VisNumBench: Evaluating Number Sense of Multimodal Large Language ModelsTengjin Weng, Jingyi Wang, Wenhao Jiang, Zhong MingICCV 2025 · 被引用 1 次
- Integrating Expertise in LLMs: Crafting a Customized Nutrition Assistant with Refined Template InstructionsAnnalisa Szymanski, Brianna L. Wimer, Oghenemaro Anuyah, Heather A. Eicher-Miller 等CHI 2024 · 被引用 27 次
- MedAraBench: Large-scale Arabic Medical Question Answering Dataset and BenchmarkMouath Abu Daoud, Leen Kharouf, Omar El Hajj, Dana El Samad 等ICLR 2026 · 被引用 4 次
- DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary DomainSong Jin, Juntian Zhang, Xun Zhang, Zeying Tian 等ACL 2026 · 被引用 1 次
