Reducing Sampling Error in Batch Temporal Difference Learning
Published in arXiv (Cornell University) • Jan 1, 2020
NobleIDNI5P41W36R17S84
Authors:,,
Brahma S. Pavse
Ishan Durugkar
Josiah P. Hanna
Abstract
Temporal difference (TD) learning is one of the main foundations of modern reinforcement learning. This paper studies the use of TD(0), a canonical TD algorithm, to estimate the value function of a given policy from a batch of data. In this batch setting, we show that TD(0) may converge to an inaccu...
Finding related papers...
Discussions
(0)No comments yet
Be the first to share your thoughts!