SciCheck: Reasoning Distillation for Biomedical Claim Verification
Gabriel Pereira, Luciano Barbosa
摘要
False claims about medical information can have damaging consequences. Although LLMs have been used to verify such claims, their success depends on access to accurate content and robust reasoning capabilities. Prior research assumes access to gold evidence during inference or relies on expensive, compute-intensive scaling strategies. However, these approaches fall short when applied to real-world verification tasks, where relevant evidence is hard to obtain and efficiency is crucial. To address this, we introduce SciCheck, a novel solution that integrates web evidence retrieval with a process of reasoning distillation. Our approach fine-tunes a small language model using reasoning traces generated by an LLM. The distillation process involves a data preparation pipeline that avoids data leakage and filters reasoning traces to retain only those leading to correct answers. It also combines web and gold evidence during training to improve robustness, while evaluation is performed with web retrieval only. We performed an extensive experimental evaluation on different claim verification datasets. The results demonstrate that SciCheck outperforms competing approaches and proprietary LLMs such as Gemini 2.5 Flash in most scenarios with lower computational cost.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Combining Evidence and Reasoning for Biomedical Fact-CheckingMariano Barone, Antonio Romano, Giuseppe Riccio, Marco Postiglione 等SIGIR 2025 · 被引用 3 次
- Knowledge-Augmented Reasoning Distillation for Small Language Models in Knowledge-Intensive TasksMinki Kang, Seanie Lee, Jinheon Baek, Kenji Kawaguchi 等NeurIPS 2023 · 被引用 128 次
- A Fact-Checking Framework with Denoising Evidence Retrieval and LLM-Based Debate VerificationJun Yang, Yuhan Bai, Dandan Song, Zhijing Wu 等WWW 2026
- Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process RewardsJaehoon Yun, Jiwoong Sohn, Jungwoo Park, Hyunjae Kim 等EMNLP 2025
- Distilling LLM Agent into Small Models with Retrieval and Code ToolsMinki Kang, Jongwon Jeong, Seanie Lee, Jaewoong Cho 等NeurIPS 2025 · 被引用 51 次
