Language Models as Science Tutors
Alexis Chevalier, Jiayi Geng, Alexander Wettig, Howard Chen, Sebastian Mizera, Toni Annala, Max Jameson Aragon, Arturo Rodríguez Fanlo, Simon Frieder, Simon Machado, Akshara Prabhakar, Ellie Thieu
摘要
NLP has recently made exciting progress toward training language models (LMs) with strong scientific problem-solving skills. However, model development has not focused on real-life usecases of LMs for science, including applications in education that require processing long scientific documents. To address this, we introduce TUTOREVAL and TUTORCHAT. TUTOREVAL is a diverse question-answering benchmark consisting of questions about long chapters from STEM textbooks, written by experts. TUTOREVAL helps measure real-life usability of LMs as scientific assistants, and it is the first benchmark combining long contexts, free-form generation, and multidisciplinary scientific knowledge. Moreover, we show that fine-tuning base models with existing dialogue datasets leads to poor performance on TUTOREVAL. Therefore, we create TUTOR-CHAT, a dataset of 80,000 long synthetic dialogues about textbooks. We use TUTORCHAT to fine-tune Llemma models with 7B and 34B parameters. These LM tutors specialized in math have a 32K-token context window, and they excel at TUTOREVAL while performing strongly on GSM8K and MATH. Our datasets build on opensource materials, and we release our models, data, and evaluations publicly. Figure 1: Example from TUTOREVAL. Given the chapter, the student asks a question to the LM Tutor. Both the chapter and the question are fed to the LM Tutor to generate the answer. GPT-4 assesses the generation by referencing the human annotated key points (blue: the tutoring task; yellow: evaluation). See detailed examples in §A.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- EXP-Bench: Can AI Conduct AI Research Experiments?Patrick Tser Jern Kon, Qiuyi Ding, Jiachen Liu, Xinyi Zhu 等ICLR 2026 · 被引用 35 次
- The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical ProofsJasper Dekoninck, Ivo Petrov, Kristian Minchev, Miroslav Marinov 等ICLR 2026 · 被引用 28 次
- When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI CollaborationQuan Shi, Carlos E. Jimenez, Shunyu Yao, Nick Haber 等NeurIPS 2025 · 被引用 5 次
- Closing the Loop: Learning to Generate Writing Feedback via Language Model Simulated Student RevisionsInderjeet Nair, Jiaye Tan, Xiaotian Su, Anne Gere 等EMNLP 2024 · 被引用 2 次
- DDxTutor: Clinical Reasoning Tutoring System with Differential Diagnosis-Based Structured ReasoningQian Wu, Zheyao Gao, Longfei Gou, Qi DouACL 2025 · 被引用 2 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
相关 Paper
- Enhancing Chat Language Models by Scaling High-quality Instructional ConversationsNing Ding, Yulin Chen, Bokai Xu, Yujia Qin 等EMNLP 2023 · 被引用 95 次
- SciAgent: Tool-augmented Language Models for Scientific ReasoningYubo Ma, Zhibin Gou, Junheng Hao, Ruochen Xu 等EMNLP 2024 · 被引用 13 次
- SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie 等EMNLP 2025
- ınftyBench: Extending Long Context Evaluation Beyond 100K TokensXinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu 等ACL 2024
- Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science CommunicatorsPrasoon Bajpai, Niladri Chatterjee, Subhabrata Dutta, Tanmoy ChakrabortyEMNLP 2024
