Learning to Generate Explainable Stock Predictions using Self-Reflective Large Language Models
Kelvin J. L. Koa, Yunshan Ma, Ritchie Ng, Tat-Seng Chua
Abstract
Explaining stock predictions is generally a difficult task for traditional non-generative deep learning models, where explanations are limited to visualizing the attention weights on important texts. Today, Large Language Models (LLMs) present a solution to this problem, given their known capabilities to generate human-readable explanations for their decision-making process. However, the task of stock prediction remains challenging for LLMs, as it requires the ability to weigh the varying impacts of chaotic social texts on stock prices. The problem gets progressively harder with the introduction of the explanation component, which requires LLMs to explain verbally why certain factors are more important than the others. On the other hand, to fine-tune LLMs for such a task, one would need expert-annotated samples of explanation for every stock movement in the training set, which is expensive and impractical to scale. To tackle these issues, we propose our Summarize-Explain-Predict (SEP) framework, which utilizes a verbal self-reflective agent and Proximal Policy Optimization (PPO) that allow a LLM teach itself how to generate explainable stock predictions, in a fully autonomous manner. The reflective agent learns how to explain past stock movements through a self-reasoning process, while the PPO trainer trains the model to generate the most likely explanations given the input texts at test-time. The training samples for the PPO trainer are also the responses generated during the reflective process, which eliminates the need for human annotators. Using our SEP framework, we fine-tune a specialized LLM that can outperform both traditional deep-learning and LLM methods in prediction accuracy and Matthews correlation coefficient, for the stock classification task. To justify the generalization capability of our framework, we further test it on the portfolio construction task, and demonstrate its effectiveness through various portfolio metrics. Our code can be accessed through https://github.com/koa-fin/sep .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53ed4879-0755-4d9f-b2d4-bf17b64a125bCited by top-tier papers10
- AURORA: Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse VerificationXiaoyu Tan, Tianchu Yao, Chao Qu, Bin Li et al.KDD 2026 · 18 citations
- EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product AssociationWeiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag et al.ACL 2025 · 15 citations
- Trade in Minutes! Rationality-Driven Agentic System for Quantitative Financial TradingZifan Song, Kaitao Song, Guosheng Hu, Ding Qi et al.ICLR 2026 · 7 citations
- Translate Policy to Language: Flow Matching Generated Rewards for LLM ExplanationsXinyi Yang, Liang Zeng, Heng Dong, Chao Yu et al.ICLR 2026 · 6 citations
- Reasoning on Time-Series for Financial Technical AnalysisKelvin J. L. Koa, Jan Chen, Yunshan Ma, Huanhuan Zheng et al.ICLR 2026 · 5 citations
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
Related papers
- Reflective Multi-Agent Collaboration based on Large Language ModelsXiaohe Bo, Zeyu Zhang, Quanyu Dai, Xueyang Feng et al.NeurIPS 2024 · 87 citations
- FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision MakingYangyang Yu, Zhiyuan Yao, Haohang Li, Zhiyang Deng et al.NeurIPS 2024 · 197 citations
- LangTime: A Language-Guided Unified Model for Time Series Forecasting with Proximal Policy OptimizationWenzhe Niu, Zongxia Xie, Yanru Sun, Wei He et al.ICML 2025
- VinePPO: Refining Credit Assignment in RL Training of LLMsAmirhossein Kazemnejad, Milad Aghajohari, Eva Portelance, Alessandro Sordoni et al.ICML 2025
- Re2LLM: Reflective Reinforcement Large Language Model for Session-based RecommendationZiyan Wang, Yingpeng Du, Zhu Sun, Haoyan Chua et al.AAAI 2025 · 10 citations
