Quadratic Exponential Decrease Roll-Back: An Efficient Gradient Update Mechanism in Proximal Policy Optimization
Published • Jul 25, 2023
NobleIDNI5P28W79R14S26
Authors:,,
Runhao Zhao
Dan Xu
Shuoqu Jian
Abstract
Proximal Policy Optimization (PPO) achieved the best performance in many continuous control tasks as a popular Deep reinforcement learning algorithm. However, the maximum point of the surrogate in the PPO and its variant version is not the zero point of the policy gradient. In this paper, we propose...
Finding related papers...
Discussions
(0)No comments yet
Be the first to share your thoughts!