Learning to Generate Explainable Stock Predictions using Self-Reflective Large Language Models
Kelvin J. L. Koa, Yunshan Ma, Ritchie Ng, Tat-Seng Chua
摘要
Explaining stock predictions is generally a difficult task for traditional non-generative deep learning models, where explanations are limited to visualizing the attention weights on important texts. Today, Large Language Models (LLMs) present a solution to this problem, given their known capabilities to generate human-readable explanations for their decision-making process. However, the task of stock prediction remains challenging for LLMs, as it requires the ability to weigh the varying impacts of chaotic social texts on stock prices. The problem gets progressively harder with the introduction of the explanation component, which requires LLMs to explain verbally why certain factors are more important than the others. On the other hand, to fine-tune LLMs for such a task, one would need expert-annotated samples of explanation for every stock movement in the training set, which is expensive and impractical to scale. To tackle these issues, we propose our Summarize-Explain-Predict (SEP) framework, which utilizes a verbal self-reflective agent and Proximal Policy Optimization (PPO) that allow a LLM teach itself how to generate explainable stock predictions, in a fully autonomous manner. The reflective agent learns how to explain past stock movements through a self-reasoning process, while the PPO trainer trains the model to generate the most likely explanations given the input texts at test-time. The training samples for the PPO trainer are also the responses generated during the reflective process, which eliminates the need for human annotators. Using our SEP framework, we fine-tune a specialized LLM that can outperform both traditional deep-learning and LLM methods in prediction accuracy and Matthews correlation coefficient, for the stock classification task. To justify the generalization capability of our framework, we further test it on the portfolio construction task, and demonstrate its effectiveness through various portfolio metrics. Our code can be accessed through https://github.com/koa-fin/sep .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- AURORA: Automated Training Framework of Universal Process Reward Models via Ensemble Prompting and Reverse VerificationXiaoyu Tan, Tianchu Yao, Chao Qu, Bin Li 等KDD 2026 · 被引用 18 次
- EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product AssociationWeiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag 等ACL 2025 · 被引用 15 次
- Trade in Minutes! Rationality-Driven Agentic System for Quantitative Financial TradingZifan Song, Kaitao Song, Guosheng Hu, Ding Qi 等ICLR 2026 · 被引用 7 次
- Translate Policy to Language: Flow Matching Generated Rewards for LLM ExplanationsXinyi Yang, Liang Zeng, Heng Dong, Chao Yu 等ICLR 2026 · 被引用 6 次
- Reasoning on Time-Series for Financial Technical AnalysisKelvin J. L. Koa, Jan Chen, Yunshan Ma, Huanhuan Zheng 等ICLR 2026 · 被引用 5 次
它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
相关 Paper
- Reflective Multi-Agent Collaboration based on Large Language ModelsXiaohe Bo, Zeyu Zhang, Quanyu Dai, Xueyang Feng 等NeurIPS 2024 · 被引用 87 次
- FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision MakingYangyang Yu, Zhiyuan Yao, Haohang Li, Zhiyang Deng 等NeurIPS 2024 · 被引用 197 次
- LangTime: A Language-Guided Unified Model for Time Series Forecasting with Proximal Policy OptimizationWenzhe Niu, Zongxia Xie, Yanru Sun, Wei He 等ICML 2025
- VinePPO: Refining Credit Assignment in RL Training of LLMsAmirhossein Kazemnejad, Milad Aghajohari, Eva Portelance, Alessandro Sordoni 等ICML 2025
- Re2LLM: Reflective Reinforcement Large Language Model for Session-based RecommendationZiyan Wang, Yingpeng Du, Zhu Sun, Haoyan Chua 等AAAI 2025 · 被引用 10 次
