SocraticChem: Physics-Grounded Socratic Inquiry for Safety-Critical Experimental Science
Jianhang Ye, Wanqi Yang, Bo Zhang, Zekun Li, Jian Zhang, Ming Yang, Yinghuan Shi
Abstract
Large language models (LLMs) have emerged as a foundational technology for intelligent education. However, in safety-critical domains like chemistry experiments, current general models face a fundamental pedagogical paradox : effective inquiry requires students to learn from errors, but the physical world imposes strict safety constraints where certain errors are impermissible ( e.g., mixing reagents incorrectly could cause explosions). Moreover, existing LLM-based tutors typically prioritize textual plausibility over physical reality, leading to ''Hallucinated Pedagogy'' when applied to real-world experiments. To address this paradox, we propose SocraticChem, a physics-grounded framework that formalizes tutoring not as open-ended generation, but as a verifiable, safety-aware decision policy. SocraticChem guides students toward learning objectives while strictly intercepting potentially dangerous actions before they manifest physically, enabling error-driven learning without physical risk. To instantiate this framework, we first develop a multi-agent LLM pipeline to construct SoChemDataset, comprising 15.2K physically grounded teaching turns across 119 middle school chemistry experiments, and then fine-tune a SoChem-LLM on this dataset. Finally, we establish a comprehensive evaluation suite spanning physics-grounded verification, LLM-based assessment, and standard NLP benchmarks. Extensive experiments demonstrate that SoChem-LLM significantly outperforms baselines, achieving a State Awareness of 70.32% (surpassing GPT-4o's 49.22%) and a Safety Score of 9.00 (surpassing the best baseline of 8.00). These results confirm its capability to strictly enforce physical safety constraints while maintaining high-quality pedagogical guidance. Our code is available at: https://github.com/bsw-ili/socChem_final.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 99d9df7d-93a5-4539-bef4-6866b3f3a329Related papers
- SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMsSihang Zhao, Kangrui Yu, Youliang Yuan, Pinjia He et al.ACL 2026
- SocraticLM: Exploring Socratic Personalized Teaching with Large Language ModelsJiayu Liu, Zhenya Huang, Tong Xiao, Jing Sha et al.NeurIPS 2024 · 65 citations
- SoSBench: Benchmarking Safety Alignment on Six Scientific DomainsFengqing Jiang, Fengbo Ma, Zhangchen Xu, Yuetai Li et al.ICLR 2026 · 14 citations
- From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement LearningDavid Dinucu-Jianu, Jakub Macina, Nico Daheim, Ido Hakimi et al.EMNLP 2025
- Grounding Generative Planners in Verifiable Logic: A Hybrid Architecture for Trustworthy Embodied AIFeiyu Wu, Xu Zheng, Yue Qu, Zhuocheng Wang et al.ICLR 2026 · 4 citations
