Reliable Evaluation and Benchmarks for Statement Autoformalization
Auguste Poiroux, Gail Weiss, Viktor Kuncak, Antoine Bosselut
摘要
Evaluating statement autoformalization, translating natural language mathematics into formal languages like Lean 4, remains a significant challenge, with few metrics, datasets, and standards to robustly measure progress. In this work, we present a comprehensive approach combining improved metrics, robust benchmarks, and systematic evaluation, to fill this gap. First, we introduce BEq+, an automated metric that correlates strongly with human judgment, along with ProofNetVerif, a new dataset for assessing the quality of evaluation metrics, containing 3,752 annotated examples. Second, we develop two new autoformalization benchmarks: ProofNet#, a corrected version of ProofNet, and RLM25, with 619 new pairs of research-level mathematics from six formalization projects. Through systematic experimentation across these benchmarks, we find that current techniques can achieve up to 45.1% accuracy on undergraduate mathematics but struggle with research-level content without proper context. Our work establishes a reliable foundation for evaluating and advancing autoformalization systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ProofBridge: Auto-Formalization of Natural Language Proofs in Lean via Joint EmbeddingsPrithwish Jana, Kaan Kale, Ahmet Ege Tanriverdi, Cruise Song 等ICLR 2026 · 被引用 19 次
- A Minimal Agent for Automated Theorem ProvingBorja Requena, Austin Letson, Krystian Nowakowski, Izan Beltran Ferreiro 等ICML 2026 · 被引用 9 次
- ASSESS: A Semantic and Structural Evaluation Framework for Statement SimilarityXiaoyang Liu, Tao Zhu, Zineng Dong, Yuntian Liu 等ICLR 2026 · 被引用 9 次
- FormalScience: Scalable Human-in-the-Loop Autoformalisation of Science with Agentic Code Generation in LeanJordan Meadows, Lan Zhang, André FreitasACL 2026 · 被引用 3 次
- Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem ProvingPawan Sasanka Ammanamanchi, Siddharth Bhat, Stella BidermanICML 2026 · 被引用 2 次
它引用的顶会 Paper9
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Llemma: An Open Language Model for MathematicsZhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos 等ICLR 2024 · 被引用 433 次
- Autoformalization with Large Language ModelsYuhuai Wu, Albert Qiaochu Jiang, Wenda Li, Markus N. Rabe 等NeurIPS 2022 · 被引用 364 次
相关 Paper
- Rethinking and Improving Autoformalization: Towards a Faithful Metric and a Dependency Retrieval-based ApproachQi Liu, Xinhao Zheng, Xudong Lu, Qinxiang Cao 等ICLR 2025
- miniF2F-Lean Revisited: Reviewing Limitations and Charting a Path ForwardAzim Ospanov, Farzan Farnia, Roozbeh MohitNeurIPS 2025 · 被引用 14 次
- CriticLean: Critic-Guided Reinforcement Learning for Mathematical FormalizationZhongyuan Peng, Yifan Yao, Kaijing Ma, Shuyue Guo 等ACL 2026 · 被引用 15 次
- ReForm: Reflective Autoformalization with Prospective Bounded Sequence OptimizationGuoxin Chen, Jing Wu, Xinjie Chen, Xin Zhao 等ICLR 2026 · 被引用 22 次
- Multi-language Diversity Benefits AutoformalizationAlbert Q. Jiang, Wenda Li, Mateja JamnikNeurIPS 2024 · 被引用 12 次
