FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction
Dong Shu, Yanguang Liu, Huopu Zhang, Mengnan Du
摘要
Predicting corporate earnings surprises is a profitable yet challenging task, as accurate forecasts can inform significant investment decisions. However, progress in this domain has been constrained by a reliance on expensive, proprietary, and text-only data, limiting the development of advanced models. To address this gap, we introduce FinCall-Surprise (Financial Conference Call for Earning Surprise Prediction), the first large-scale, open-source, and multi-modal dataset for earnings surprise prediction. Comprising 2,688 unique corporate conference calls from 2019 to 2021, our dataset features word-to-word conference call textual transcripts, full audio recordings, and corresponding presentation slides. We establish a comprehensive benchmark by evaluating 26 state-of-the-art unimodal and multi-modal LLMs. Our findings reveal that (1) while many models achieve high accuracy, this performance is often an illusion caused by significant class imbalance in the real-world data. (2) Some specialized financial models demonstrate unexpected weaknesses in instruction-following and language generation. (3) Although incorporating audio and visual modalities provides some performance gains, current models still struggle to leverage these signals effectively. These results highlight critical limitations in the financial reasoning capabilities of existing LLMs and establish a challenging new baseline for future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Instruction Pre-Training: Language Models are Supervised Multitask LearnersDaixuan Cheng, Yuxian Gu, Shaohan Huang, Junyu Bi 等EMNLP 2024 · 被引用 13 次
- FinChart-Bench: Benchmarking Financial Chart Comprehension in Vision-Language ModelsDong Shu, Haoyang Yuan, Yuchen Wang, Yanguang Liu 等ACL 2026 · 被引用 11 次
- FinTextQA: A Dataset for Long-form Financial Question AnsweringJian Chen, Peilin Zhou, Yining Hua, Loh Xin 等ACL 2024 · 被引用 7 次
- RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented InstructionsWanlong Liu, Junying Chen, Ke Ji, Li Zhou 等EMNLP 2025 · 被引用 1 次
相关 Paper
- NumHTML: Numeric-Oriented Hierarchical Transformer Model for Multi-Task Financial ForecastingLinyi Yang, Jiazheng Li, Ruihai Dong, Yue Zhang 等AAAI 2022 · 被引用 54 次
- Multimodal Multi-Speaker Merger & Acquisition Financial Modeling: A New Task, Dataset, and Neural BaselinesRamit Sawhney, Mihir Goyal, Prakhar Goel, Puneet Mathur 等ACL 2021
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and ChallengingZichen Tang, Haihong E, Jiacheng Liu, Zhongjun Yang 等ICCV 2025 · 被引用 1 次
- MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial ApplicationXueqing Peng, Lingfei Qian, Yan Wang, Ruoyu Xiang 等ACL 2026 · 被引用 6 次
- EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial StatementsIssa Sugiura, Takashi Ishida, Taro Makino, Chieko Tazuke 等ICLR 2026 · 被引用 9 次
