Measuring Consistency in Text-based Financial Forecasting Models
Linyi Yang, Yingpeng Ma, Yue Zhang
Abstract
Financial forecasting has been an important and active area of machine learning research, as even the most modest advantage in predictive accuracy can be parlayed into significant financial gains. Recent advances in natural language processing (NLP) bring the opportunity to leverage textual data, such as earnings reports of publicly traded companies, to predict the return rate for an asset. However, when dealing with such a sensitive task, the consistency of models -their invariance under meaning-preserving alternations in input -is a crucial property for building user trust. Despite this, current financial forecasting methods do not consider consistency. To address this problem, we propose FinTrust, an evaluation tool that assesses logical consistency in financial text. Using FinTrust, we show that the consistency of state-of-the-art NLP models for financial forecasting is poor. Our analysis of the performance degradation caused by meaning-preserving alternations suggests that current text-based methods are not suitable for robustly predicting market information. All resources are available at https: //github.com/yingpengma/FinTrust .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Learning The Difference That Makes A Difference With Counterfactually-Augmented DataDivyansh Kaushik, Eduard H. Hovy, Zachary Chase LiptonICLR 2020 · 625 citations
- HTML: Hierarchical Transformer-based Multi-task Learning for Volatility PredictionLinyi Yang, Tin Lok James Ng, Barry Smyth, Ruihai DongWWW 2020 · 127 citations
- Robustness to Spurious Correlations via Human AnnotationsMegha Srivastava, Tatsunori B. Hashimoto, Percy LiangICML 2020 · 103 citations
- An Analysis of Natural Language Inference Benchmarks through the Lens of NegationMd Mosharaf Hossain, Venelin Kovatchev, Pranoy Dutta, Tiffany Kao et al.EMNLP 2020 · 61 citations
Related papers
- FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance DomainTiansheng Hu, Tongyan Hu, Liuyang Bai, Yilun Zhao et al.EMNLP 2025 · 2 citations
- NumHTML: Numeric-Oriented Hierarchical Transformer Model for Multi-Task Financial ForecastingLinyi Yang, Jiazheng Li, Ruihai Dong, Yue Zhang et al.AAAI 2022 · 54 citations
- Transitive self-consistency evaluation of NLI models without gold labelsWei Wu, Mark LastEMNLP 2025
- Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?Ashutosh Bajpai, Tanmoy ChakrabortyEMNLP 2025
- EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial StatementsIssa Sugiura, Takashi Ishida, Taro Makino, Chieko Tazuke et al.ICLR 2026 · 9 citations
