Lune

ACL2026顶会

SciCoQA: Quality Assurance for Scientific Paper-Code Alignment

Tim Baumgärtner, Iryna Gurevych

2026年份
5被引次数

摘要

Discrepancies between scientific papers and their code undermine reproducibility, a concern that grows as automated research agents scale scientific output beyond human review capacity. Whether LLMs can reliably detect such discrepancies has not been systematically measured. To this end, we present SCICOQA, a dataset of 635 paper-code discrepancies (92 real, 543 synthetic) for this cross-modal verification task. Across 22 evaluated models, even the best-performing LLMs, Gemini 3.1 Pro and GPT-5 Mini, detect only 46.7% of real-world discrepancies, revealing a critical gap in automated scientific quality assurance. We construct SCICOQA from GitHub issues and reproducibility papers, and propose a synthetic generation pipeline to scale beyond AI to Physics, Quantitative Biology, and other computational sciences. We further introduce a taxonomy of discrepancy types and categories to characterize the occurring mismatches. Our analysis shows that models particularly struggle with omitted paper details, long-context inputs, and papers outside their pre-training corpus. 1 We feed λ to two MLPs M σ and M μ to generate two, σ and μ, of dimensionality C each. We then multiply the feature map channel-wise by σ and add μ to get the transformed feature map: f ̃ijk = σ k f ijk + μ k , σ = M σ (λ), μ = M μ (λ) def forward(self, x): m = self.mu(x) s = self.sigma(x) return F.sigmoid(m), F.sigmoid(s)

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖