Universal Trading for Order Execution with Oracle Policy Distillation
Yuchen Fang, Kan Ren, Weiqing Liu, Dong Zhou, Weinan Zhang, Jiang Bian, Yong Yu, Tie-Yan Liu
摘要
As a fundamental problem in algorithmic trading, order execution aims at fulfilling a specific trading order, either liquidation or acquirement, for a given instrument. Towards effective execution strategy, recent years have witnessed the shift from the analytical view with model-based market assumptions to model-free perspective, i.e., reinforcement learning, due to its nature of sequential decision optimization. However, the noisy and yet imperfect market information that can be leveraged by the policy has made it quite challenging to build up sample efficient reinforcement learning methods to achieve effective order execution. In this paper, we propose a novel universal trading policy optimization framework to bridge the gap between the noisy yet imperfect market states and the optimal action sequences for order execution. Particularly, this framework leverages a policy distillation method that can better guide the learning of the common policy towards practically optimal execution by an oracle teacher with perfect information to approximate the optimal trading strategy. The extensive experiments have shown significant improvements of our method over various strong baselines, with reasonable trading actions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Bootstrapped Transformer for Offline Reinforcement LearningKerong Wang, Hanye Zhao, Xufang Luo, Kan Ren 等NeurIPS 2022 · 被引用 54 次
- PerfectDou: Dominating DouDizhu with Perfect Information DistillationGuan Yang, Minghuan Liu, Weijun Hong, Weinan Zhang 等NeurIPS 2022 · 被引用 41 次
- Hindsight Learning for MDPs with Exogenous InputsSean R. Sinclair, Felipe Vieira Frujeri, Ching-An Cheng, Luke Marshall 等ICML 2023 · 被引用 31 次
- Reinforcement Learning with Automated Auxiliary Loss SearchTairan He, Yuge Zhang, Kan Ren, Minghuan Liu 等NeurIPS 2022 · 被引用 22 次
- Reinforcement Learning with Maskable Stock Representation for Portfolio Management in Customizable Stock PoolsWentao Zhang, Yilei Zhao, Shuo Sun, Jie Ying 等WWW 2024 · 被引用 18 次
它引用的顶会 Paper2
相关 Paper
- FreeKD: Free-direction Knowledge Distillation for Graph Neural NetworksKaituo Feng, Changsheng Li, Ye Yuan, Guoren WangKDD 2022 · 被引用 28 次
- In-context Reinforcement Learning with Algorithm DistillationMichael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto 等ICLR 2023 · 被引用 10 次
- MetaTrader: Learning to Generalize RL Trading Policies Beyond Offline DataHaochen Yuan, Minting Pan, Yunbo Wang, Siyu Gao 等AAAI 2026
- Ranking Policy GradientKaixiang Lin, Jiayu ZhouICLR 2020 · 被引用 8 次
- Provable Partially Observable Reinforcement Learning with Privileged InformationYang Cai, Xiangyu Liu, Argyris Oikonomou, Kaiqing ZhangNeurIPS 2024 · 被引用 22 次
