athenara:~$ registry cite papers/opd-order-execution
Universal Trading for Order Execution with Oracle Policy Distillation
Introduces Oracle Policy Distillation for reinforcement-learning order execution; the resulting OPDS method ships as a runnable workflow in Microsoft Qlib.
added 2026-08-17 · MIT · external
$ git clone https://github.com/microsoft/qlib/tree/main/examples/rl_order_execution@article{opd-order-execution, title = {Universal Trading for Order Execution with Oracle Policy Distillation}, author = {Yuchen Fang and Kan Ren and Weiqing Liu and Dong Zhou and Weinan Zhang and Jiang Bian and Yong Yu and Tie-Yan Liu and Microsoft Research and Shanghai Jiao Tong University}, year = {2021}, eprint = {2103.10860}, archiveprefix = {arXiv}, note = {AAAI 2021}, }
A universal trading policy optimization framework for order execution — deciding how to slice and time the trades that fill an order. Its contribution is policy distillation: a teacher policy with perfect information guides the learning of a common policy that only ever sees noisy, imperfect market states. Published in the Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35 no. 1, pp. 107–115 (doi:10.1609/aaai.v35i1.16083).
Microsoft Qlib ships the method as OPDS under examples/rl_order_execution, whose README names it
as the AAAI 2021 paper’s method and pairs it with a PPO baseline and a TWAP weak baseline; the
directory carries train_opds.yml and backtest_opds.yml configs driving
qlib.rl.contrib.train_onpolicy and qlib.rl.contrib.backtest. This is live code rather than an
archived artifact — the example was last touched in March 2026 inside a repository whose latest
commit is July 2026, under Qlib’s MIT license.
Two caveats before starting. No data is bundled: the workflow expects you to build 5-minute China
A-share (hs300) Qlib data yourself before training. And two code locations exist and are not
equivalent — the authors’ project page points at the older
high-freq-execution branch, frozen since November 2022, while the maintained implementation is
the rl_order_execution example on main. The improvements the paper reports over its baselines
are the authors’ own experimental results on their own data.
trading [●●●●·] advanced ai [●●●●·] advanced programming [●●●··] moderate setup [●●●●·] advanced
athenara:~$