
A Unifying Perspective of Parametric Policy Search Methods for Markov Decision Processes
课程网址: http://videolectures.net/nips2012_furmston_processes/  
主讲教师: Thomas Furmston
开课单位: 伦敦大学学院
开课时间: 2013-01-16
课程语种: 英语
课程简介: Parametric policy search algorithms are one of the methods of choice for the optimisation of Markov Decision Processes, with Expectation Maximisation and natural gradient ascent being considered the current state of the art in the field. In this article we provide a unifying perspective of these two algorithms by showing that their step-directions in the parameter space are closely related to the search direction of an approximate Newton method. This analysis leads naturally to the consideration of this approximate Newton method as an alternative gradient-based method for Markov Decision Processes. We are able show that the algorithm has numerous desirable properties, absent in the naive application of Newton's method, that make it a viable alternative to either Expectation Maximisation or natural gradient ascent. Empirical results suggest that the algorithm has excellent convergence and robustness properties, performing strongly in comparison to both Expectation Maximisation and natural gradient ascent.
关 键 词: 参数策略搜索算法; 马尔可夫决策; 牛顿法; 收敛性; 鲁棒性
课程来源: 视频讲座网
最后编审: 2020-10-22:chenxin
阅读次数: 55