FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction
Dong Shu, Yanguang Liu, Huopu Zhang, Mengnan Du
Abstract
Predicting corporate earnings surprises is a profitable yet challenging task, as accurate forecasts can inform significant investment decisions. However, progress in this domain has been constrained by a reliance on expensive, proprietary, and text-only data, limiting the development of advanced models. To address this gap, we introduce FinCall-Surprise (Financial Conference Call for Earning Surprise Prediction), the first large-scale, open-source, and multi-modal dataset for earnings surprise prediction. Comprising 2,688 unique corporate conference calls from 2019 to 2021, our dataset features word-to-word conference call textual transcripts, full audio recordings, and corresponding presentation slides. We establish a comprehensive benchmark by evaluating 26 state-of-the-art unimodal and multi-modal LLMs. Our findings reveal that (1) while many models achieve high accuracy, this performance is often an illusion caused by significant class imbalance in the real-world data. (2) Some specialized financial models demonstrate unexpected weaknesses in instruction-following and language generation. (3) Although incorporating audio and visual modalities provides some performance gains, current models still struggle to leverage these signals effectively. These results highlight critical limitations in the financial reasoning capabilities of existing LLMs and establish a challenging new baseline for future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 290c8aa8-733d-4861-946f-2b76e8fc1389Builds on7
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Instruction Pre-Training: Language Models are Supervised Multitask LearnersDaixuan Cheng, Yuxian Gu, Shaohan Huang, Junyu Bi et al.EMNLP 2024 · 13 citations
- FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language ModelsDong Shu, Haoyang Yuan, Yuchen Wang, Yanguang Liu et al.ACL 2026 · 11 citations
- FinTextQA: A Dataset for Long-form Financial Question AnsweringJian Chen, Peilin Zhou, Yining Hua, Loh Xin et al.ACL 2024 · 7 citations
- RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented InstructionsWanlong Liu, Junying Chen, Ke Ji, Li Zhou et al.EMNLP 2025 · 1 citation
Related papers
- NumHTML: Numeric-Oriented Hierarchical Transformer Model for Multi-Task Financial ForecastingLinyi Yang, Jiazheng Li, Ruihai Dong, Yue Zhang et al.AAAI 2022 · 54 citations
- Multimodal Multi-Speaker Merger & Acquisition Financial Modeling: A New Task, Dataset, and Neural BaselinesRamit Sawhney, Mihir Goyal, Prakhar Goel, Puneet Mathur et al.ACL 2021
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and ChallengingZichen Tang, Haihong E, Jiacheng Liu, Zhongjun Yang et al.ICCV 2025 · 1 citation
- MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial ApplicationXueqing Peng, Lingfei Qian, Yan Wang, Ruoyu Xiang et al.ACL 2026 · 6 citations
- EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial StatementsIssa Sugiura, Takashi Ishida, Taro Makino, Chieko Tazuke et al.ICLR 2026 · 9 citations
