FormalRx: Rectify and eXamine Semantic Failures in Autoformalization
Haocheng Wang, Baiyu Huang, Yingjia Wan, Xiao Zhu, Xiaoyang Liu, Yinya Huang, Zhijiang Guo
摘要
The veracious semantic alignment in autoformalization is significant for formal mathematical reasoning. However, existing evaluations provide only opaque binary verdicts or scalar scores, offering no interpretable insight into where or why translations fail. This opacity severely limits both human understanding and automated system improvement. To bridge this gap, we introduce FormalRx, a comprehensive diagnostic evaluation framework that transforms autoformalization assessment from black-box judgments into actionable feedback. At its core is SCI Error Taxonomy, a hierarchical classification scheme decomposing autoformalization errors into 28 distinct categories with strict priority ordering. Building on this taxonomy, FormalRx provides four critical diagnostic capabilities: alignment verdicts, error categorization, error localization, and correction. We instantiate the framework with a diagnostic model FormalRx-8B, trained on 56,287 synthetically generated samples with fine-grained diagnostic annotations, and release FormalRx-Test as the first fine-grained diagnostic benchmark. FormalRx-8B achieves F1-scores of 0.88 (verdict) and 0.71 (categorization), along with accuracies of 0.75 (localization) and 0.73 (correction), substantially outperforming both general-purpose LLMs and specialized baselines. By connecting evaluation with actionable insights, FormalRx enables systematic diagnosis and improvement of autoformalization systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Autoformalization with Large Language ModelsYuhuai Wu, Albert Qiaochu Jiang, Wenda Li, Markus N. Rabe 等NeurIPS 2022 · 被引用 364 次
- miniF2F: a cross-system benchmark for formal Olympiad-level mathematicsKunhao Zheng, Jesse Michael Han, Stanislas PoluICLR 2022 · 被引用 342 次
- ReForm: Reflective Autoformalization with Prospective Bounded Sequence OptimizationGuoxin Chen, Jing Wu, Xinjie Chen, Xin Zhao 等ICLR 2026 · 被引用 22 次
- Aria: an Agent for Retrieval and Iterative Auto-Formalization via Dependency GraphHanyu Wang, Ruohan Xie, Yutong Wang, Guoxiong Gao 等ICLR 2026 · 被引用 21 次
相关 Paper
- MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and CorrectionYuyang Wu, Jinhui Ye, Shuhao Zhang, Lu Dai 等EMNLP 2025 · 被引用 1 次
- SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic VerificationWenyao Cui, Huaping Zhang, Yongyi Huang, Qiuchi Li 等KDD 2026
- FormalAlign: Automated Alignment Evaluation for AutoformalizationJianqiao Lu, Yingjia Wan, Yinya Huang, Jing Xiong 等ICLR 2025
- PRISM: Probing Reasoning, Instruction, and Source Memory in LLM HallucinationsYuhe Wu, Guangyu Wang, Yuran Chen, Jiatong Zhang 等ACL 2026
- StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs Through Knowledge-Reasoning FusionYutong Wu, Di Huang, Ruosi Wan, Yue Peng 等AAAI 2026 · 被引用 10 次
