NobleBlocks
Public

Quadratic Exponential Decrease Roll-Back: An Efficient Gradient Update Mechanism in Proximal Policy Optimization

Published • Jul 25, 2023
NobleIDNI5P28W79R14S26
Authors:
Runhao Zhao
,
Dan Xu
,
Shuoqu Jian

Abstract

Proximal Policy Optimization (PPO) achieved the best performance in many continuous control tasks as a popular Deep reinforcement learning algorithm. However, the maximum point of the surrogate in the PPO and its variant version is not the zero point of the policy gradient. In this paper, we propose...

Finding related papers...

Discussions

(0)

No comments yet

Be the first to share your thoughts!