AlphaBench: Benchmarking Large Language Models in Formulaic Alpha Factor Mining
Haochen Luo, Ho Tin Ko, Jiandong Chen, David Q. Sun, Yuan Zhang, Chen Liu
Abstract
Formulaic alpha factor mining (FAFM) is a central problem in quantitative investment, where interpretable formulas are designed to extract predictive signals from historical financial series. With the emergence of large language models (LLMs), recent studies have begun to explore their roles in FAFM, yet their capabilities across different tasks and configurations remain unclear. In this work, we introduce AlphaBench, the first systematic benchmark for evaluating LLMs in FAFM. AlphaBench covers three core tasks, including factor generation, factor evaluation, and factor searching, which are all popular tasks integrated in the workflow of quantitative researchers. Beyond task-level evaluation, we further analyze how different LLM settings, including model type, prompting paradigm, and reasoning strategy, influence performance. Our experiments on a range of open-source and closed-source models reveal that LLMs hold strong potential in automating factor mining, while also facing persistent challenges in robustness, search efficiency, and practical usability. The project is available at: https://alphabench.cc/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47f3677c-1ccc-4bf1-b44a-5404228c642bCited by top-tier papers1
Ask how each one uses itBuilds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun et al.VLDB 2024 · 609 citations
- ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question AnsweringZhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma et al.EMNLP 2022 · 57 citations
Related papers
- Navigating the Alpha Jungle: An LLM-Powered MCTS Framework for Formulaic Alpha Factor MiningYu Shi, Yitong Duan, Jian LiAAAI 2026 · 11 citations
- INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based AgentHaohang Li, Yupeng Cao, Yangyang Yu, Shashidhar Reddy Javaji et al.ACL 2025
- BizBench: A Quantitative Reasoning Benchmark for Business and FinanceMichael Krumdick, Rik Koncel-Kedziorski, Viet Dac Lai, Varshini Reddy et al.ACL 2024 · 10 citations
- AlphaAgent: LLM-Driven Alpha Mining with Regularized Exploration to Counteract Alpha DecayZiyi Tang, Zechuan Chen, Jiarui Yang, Jiayao Mai et al.KDD 2025 · 3 citations
- Cognitive Alpha Mining via LLM-Driven Code-Based EvolutionFengyuan Liu, Yi Huang, Sichun Luo, Yuqi Wang et al.ACL 2026 · 3 citations
