Reducing sampling error in batch temporal difference learning
Published in arXiv (Cornell University) • May 1, 2020
NobleIDNI6P17W42R97S86
Authors:
Brahma S. Pavse
Abstract
Temporal difference (TD) learning is one of the main foundations of modern reinforcement learning. This thesis studies the use of TD(0), a canonical TD algorithm, to estimate the value function of a given evaluation policy from a batch of data. In this batch setting, we show that TD(0) may converge ...
Finding related papers...
Discussions
(0)No comments yet
Be the first to share your thoughts!