Journal of the Operations Research Society of China ›› 2026, Vol. 14 ›› Issue (2): 364-396.doi: 10.1007/s40305-024-00546-z

• • 上一篇    下一篇

  

  • 收稿日期:2023-02-14 修回日期:2024-03-24 出版日期:2026-06-30 发布日期:2026-07-06
  • 通讯作者: Zai-Wen Wen E-mail:wenzw@pku.edu.cn
  • 作者简介:Ze-Yu Jin,E-mail:1801210096@pku.edu.cn;Johann Michael Schmitt,E-mail:MichelSchmitt2@web.de

On the Analysis of Model-Free Methods for the Linear Quadratic Regulator

Ze-Yu Jin1, Johann Michael Schmitt2, Zai-Wen Wen2   

  1. 1 School of Mathematical Sciences, Peking University, Beijing 100871, China;
    2 Beijing International Center for Mathematical Research, Peking University, Beijing 100871, China
  • Received:2023-02-14 Revised:2024-03-24 Online:2026-06-30 Published:2026-07-06
  • Contact: Zai-Wen Wen E-mail:wenzw@pku.edu.cn
  • Supported by:
    This work was supported in part by the National Natural Science Foundation of China (No.12331010).

Abstract: Many reinforcement learning methods achieve great success in practice but lack theoretical foundation. In this paper, we study the convergence analysis of the model-free methods for the Linear Quadratic Regulator by treating the underlying system as a black box. The global linear convergence properties and sample complexities are established for several popular algorithms such as the temporal differences (TD)-learning method, the policy gradient algorithm, and the actor-critic (AC) algorithm. Our analysis shows that the actor-critic algorithm can reduce the sample complexity compared with the policy gradient algorithm. Although our analysis is still preliminary, it still explains the benefit of AC algorithm in a certain sense.

Key words: Linear Quadratic Regulator, TD-learning, The policy gradient algorithm, The actor-critic algorithm, Convergence

中图分类号: