Journal of the Operations Research Society of China ›› 2026, Vol. 14 ›› Issue (2): 364-396.doi: 10.1007/s40305-024-00546-z

Previous Articles     Next Articles

On the Analysis of Model-Free Methods for the Linear Quadratic Regulator

Ze-Yu Jin1, Johann Michael Schmitt2, Zai-Wen Wen2   

  1. 1 School of Mathematical Sciences, Peking University, Beijing 100871, China;
    2 Beijing International Center for Mathematical Research, Peking University, Beijing 100871, China
  • Received:2023-02-14 Revised:2024-03-24 Online:2026-06-30 Published:2026-07-06
  • Contact: Zai-Wen Wen E-mail:wenzw@pku.edu.cn
  • Supported by:
    This work was supported in part by the National Natural Science Foundation of China (No.12331010).

Abstract: Many reinforcement learning methods achieve great success in practice but lack theoretical foundation. In this paper, we study the convergence analysis of the model-free methods for the Linear Quadratic Regulator by treating the underlying system as a black box. The global linear convergence properties and sample complexities are established for several popular algorithms such as the temporal differences (TD)-learning method, the policy gradient algorithm, and the actor-critic (AC) algorithm. Our analysis shows that the actor-critic algorithm can reduce the sample complexity compared with the policy gradient algorithm. Although our analysis is still preliminary, it still explains the benefit of AC algorithm in a certain sense.

Key words: Linear Quadratic Regulator, TD-learning, The policy gradient algorithm, The actor-critic algorithm, Convergence

CLC Number: