SocraticChem: Physics-Grounded Socratic Inquiry for Safety-Critical Experimental Science
Jianhang Ye, Wanqi Yang, Bo Zhang, Zekun Li, Jian Zhang, Ming Yang, Yinghuan Shi
摘要
Large language models (LLMs) have emerged as a foundational technology for intelligent education. However, in safety-critical domains like chemistry experiments, current general models face a fundamental pedagogical paradox : effective inquiry requires students to learn from errors, but the physical world imposes strict safety constraints where certain errors are impermissible ( e.g., mixing reagents incorrectly could cause explosions). Moreover, existing LLM-based tutors typically prioritize textual plausibility over physical reality, leading to ''Hallucinated Pedagogy'' when applied to real-world experiments. To address this paradox, we propose SocraticChem, a physics-grounded framework that formalizes tutoring not as open-ended generation, but as a verifiable, safety-aware decision policy. SocraticChem guides students toward learning objectives while strictly intercepting potentially dangerous actions before they manifest physically, enabling error-driven learning without physical risk. To instantiate this framework, we first develop a multi-agent LLM pipeline to construct SoChemDataset, comprising 15.2K physically grounded teaching turns across 119 middle school chemistry experiments, and then fine-tune a SoChem-LLM on this dataset. Finally, we establish a comprehensive evaluation suite spanning physics-grounded verification, LLM-based assessment, and standard NLP benchmarks. Extensive experiments demonstrate that SoChem-LLM significantly outperforms baselines, achieving a State Awareness of 70.32% (surpassing GPT-4o's 49.22%) and a Safety Score of 9.00 (surpassing the best baseline of 8.00). These results confirm its capability to strictly enforce physical safety constraints while maintaining high-quality pedagogical guidance. Our code is available at: https://github.com/bsw-ili/socChem_final.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMsSihang Zhao, Kangrui Yu, Youliang Yuan, Pinjia He 等ACL 2026
- SocraticLM: Exploring Socratic Personalized Teaching with Large Language ModelsJiayu Liu, Zhenya Huang, Tong Xiao, Jing Sha 等NeurIPS 2024 · 被引用 65 次
- SoSBench: Benchmarking Safety Alignment on Six Scientific DomainsFengqing Jiang, Fengbo Ma, Zhangchen Xu, Yuetai Li 等ICLR 2026 · 被引用 14 次
- From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement LearningDavid Dinucu-Jianu, Jakub Macina, Nico Daheim, Ido Hakimi 等EMNLP 2025
- Grounding Generative Planners in Verifiable Logic: A Hybrid Architecture for Trustworthy Embodied AIFeiyu Wu, Xu Zheng, Yue Qu, Zhuocheng Wang 等ICLR 2026 · 被引用 4 次
