
Separated Trust Regions Policy Optimization Method
课程网址: http://videolectures.net/kdd2019_zou_zhuang_cheng/  
主讲教师: Luobao Zou
开课单位: 上海交通大学
开课时间: 2020-03-02
课程语种: 英语


课程简介: In this work, we propose a moderate policy update method for reinforcement learning, which encourages the agent to explore more boldly in early episodes but updates the policy more cautious. Based on the maximum entropy framework, we propose a softer objective with more conservative constraints and build the separated trust regions for optimization. To reduce the variance of expected entropy return, a calculated state policy entropy of Gaussian distribution is preferred instead of collecting log probability by sampling. This new method, which we call separated trust region for policy mean and variance (STRMV), can be view as an extension to proximal policy optimization (PPO) but it is gentler for policy update and more lively for exploration. We test our approach on a wide variety of continuous control benchmark tasks in the MuJoCo environment. The experiments demonstrate that STRMV outperforms the previous state of art on-policy methods, not only achieving higher rewards but also improving the sample efficiency.
关 键 词: 最大熵框架; 信任区域; 高斯分布
课程来源: 视频讲座网
数据采集: 2020-11-29:cjy
最后编审: 2020-11-29:cjy
阅读次数: 48