SCRIBE: Structured Chain Reasoning for Interactive Behaviour Explanations using Tool Calling
Fares Fawzi, Vinitra Swamy, Dominik Glandorf, Tanya Nazaretsky, Tanja Käser
摘要
Language models can be used to provide interactive, personalized student feedback in educational settings.However, real-world deployment faces three key challenges: privacy concerns, limited computational resources, and the need for pedagogically valid responses.These constraints require small, open-source models that can run locally and reliably ground their outputs in correct information.We introduce SCRIBE, a framework for multi-hop, tool-augmented reasoning designed to generate valid responses to student questions about feedback reports.SCRIBE combines domainspecific tools with a self-reflective inference pipeline that supports iterative reasoning, tool use, and error recovery.We distil these capabilities into 3B and 8B models via two-stage LoRA fine-tuning on synthetic GPT-4o-generated data.Evaluation with a human-aligned GPT-Judge and a user study with 108 students shows that 8B-SCRIBE models achieve comparable or superior quality to much larger models in key dimensions such as relevance and actionability, while being perceived on par with GPT-4o and Llama-3.3 70B by students.These findings demonstrate the viability of SCRIBE for low-resource, privacy-sensitive educational applications.How can I improve my performance to pass the course?Useful
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 被引用 1,715 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
相关 Paper
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen 等ICLR 2024 · 被引用 308 次
- Designing Scaffolding Cards to Facilitate LLM-Based Socratic Instruction: An Exploratory Study of Response Strategies to Support LearningLujin Mao, Linyuan Dong, Wenan Li, Xiangen Hu 等CHI 2026 · 被引用 3 次
- LogicSAGE: Neuro-Symbolic Reasoning with Socratic-Guided EnhancementJinlong Tian, Jiang Yu, Kewei Cheng, Fengxiang Cheng 等ICML 2026
- GPT4Tools: Teaching Large Language Model to Use Tools via Self-instructionRui Yang, Lin Song, Yanwei Li, Sijie Zhao 等NeurIPS 2023 · 被引用 340 次
- Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive PromptingHeming Xia, Cunxiao Du, Rui Li, Chak Tou Leong 等ACL 2026
