SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs
Sihang Zhao, Kangrui Yu, Youliang Yuan, Pinjia He, Hongyi Wen
Abstract
Large Language Models (LLMs) have been widely explored in educational scenarios. We identify a critical vulnerability in current educational LLMs, pedagogical jailbreaks, where students use answer-inducing prompts to elicit solutions rather than scaffolded instructions. To enable systematic study, we unify and formalize safe, helpful, and pedagogical behaviors with a knowledge-mastery graph and introduce SHAPE, a benchmark of 9,087 student-question pairs for evaluating tutoring behavior under adversarial pressure. We propose a graph-augmented tutoring pipeline that infers prerequisite concepts from queries, identifies mastery gaps, and routes generation between instructing and problem-solving via explicit gating. Experiments across multiple LLMs show that our method yields significantly improved safety under two pedagogical jailbreak settings, while maintaining near-ceiling helpfulness under the same evaluation protocol. Our code and data are available at https://github.com/MAPS-research/SHaPE
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78fb2f98-0ffa-479c-b444-14b3b804c63bBuilds on7
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang et al.ICLR 2024 · 441 citations
- SocraticLM: Exploring Socratic Personalized Teaching with Large Language ModelsJiayu Liu, Zhenya Huang, Tong Xiao, Jing Sha et al.NeurIPS 2024 · 65 citations
- Enhancing Personalized Multi-Turn Dialogue with Curiosity RewardYanming Wan, Jiaxing Wu, Marwa Abdulhai, Lior Shani et al.NeurIPS 2025 · 32 citations
- PAIGE: Examining Learning Outcomes and Experiences with Personalized AI-Generated Educational PodcastsTiffany D. Do, Usama Bin Shafqat, Elsie Ling, Nikhil SardaCHI 2025 · 28 citations
Related papers
- Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student AttacksJin Zhao, Marta Knezevic, Tanja KäserACL 2026 · 2 citations
- MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM SafetyJialin Song, Xiaodong Liu, Weiwei Yang, Wuyang Chen et al.ICML 2026 · 5 citations
- SocraticChem: Physics-Grounded Socratic Inquiry for Safety-Critical Experimental ScienceJianhang Ye, Wanqi Yang, Bo Zhang, Zekun Li et al.KDD 2026
- GraphShield: Graph-Theoretic Modeling of Network-Level Dynamics for Robust Jailbreak DetectionSunghee Dong, Sungwon Yi, Kangmin Bae, Jaeyoon Kim et al.ICLR 2026
- EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated TeachersYilin Jiang, Mingzi Zhang, Xuanyu Yin, Sheng Jin et al.AAAI 2026 · 1 citation
