NobleBlocks
Public

Reducing Sampling Error in Batch Temporal Difference Learning

Published in arXiv (Cornell University) • Jan 1, 2020
NobleIDNI5P41W36R17S84
Authors:
Brahma S. Pavse
,
Ishan Durugkar
,
Josiah P. Hanna

Abstract

Temporal difference (TD) learning is one of the main foundations of modern reinforcement learning. This paper studies the use of TD(0), a canonical TD algorithm, to estimate the value function of a given policy from a batch of data. In this batch setting, we show that TD(0) may converge to an inaccu...

Finding related papers...

Discussions

(0)

No comments yet

Be the first to share your thoughts!