athenara:~$ registry cite papers/opd-order-execution

Universal Trading for Order Execution with Oracle Policy Distillation

Introduces Oracle Policy Distillation for reinforcement-learning order execution; the resulting OPDS method ships as a runnable workflow in Microsoft Qlib.

#order-execution #reinforcement-learning #policy-distillation #market-microstructure #qlib

added 2026-08-17 · MIT · external

$ git clone https://github.com/microsoft/qlib/tree/main/examples/rl_order_execution
@article{opd-order-execution,
  title         = {Universal Trading for Order Execution with Oracle Policy Distillation},
  author        = {Yuchen Fang and Kan Ren and Weiqing Liu and Dong Zhou and Weinan Zhang and Jiang Bian and Yong Yu and Tie-Yan Liu and Microsoft Research and Shanghai Jiao Tong University},
  year          = {2021},
  eprint        = {2103.10860},
  archiveprefix = {arXiv},
  note          = {AAAI 2021},
}

A universal trading policy optimization framework for order execution — deciding how to slice and time the trades that fill an order. Its contribution is policy distillation: a teacher policy with perfect information guides the learning of a common policy that only ever sees noisy, imperfect market states. Published in the Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35 no. 1, pp. 107–115 (doi:10.1609/aaai.v35i1.16083).

Microsoft Qlib ships the method as OPDS under examples/rl_order_execution, whose README names it as the AAAI 2021 paper’s method and pairs it with a PPO baseline and a TWAP weak baseline; the directory carries train_opds.yml and backtest_opds.yml configs driving qlib.rl.contrib.train_onpolicy and qlib.rl.contrib.backtest. This is live code rather than an archived artifact — the example was last touched in March 2026 inside a repository whose latest commit is July 2026, under Qlib’s MIT license.

Two caveats before starting. No data is bundled: the workflow expects you to build 5-minute China A-share (hs300) Qlib data yourself before training. And two code locations exist and are not equivalent — the authors’ project page points at the older high-freq-execution branch, frozen since November 2022, while the maintained implementation is the rl_order_execution example on main. The improvements the paper reports over its baselines are the authors’ own experimental results on their own data.

(END)

trading [●●●●·] advanced   ai [●●●●·] advanced   programming [●●●··] moderate   setup [●●●●·] advanced

origin external
license MIT
doi 10.1609/aaai.v35i1.16083
markets equities

implements rl-policy

builds on qlib

athenara:~$