NobleBlocks
Public

Reducing sampling error in batch temporal difference learning

Published in arXiv (Cornell University) • May 1, 2020
NobleIDNI6P17W42R97S86
Authors:
Brahma S. Pavse

Abstract

Temporal difference (TD) learning is one of the main foundations of modern reinforcement learning. This thesis studies the use of TD(0), a canonical TD algorithm, to estimate the value function of a given evaluation policy from a batch of data. In this batch setting, we show that TD(0) may converge ...

Finding related papers...

Discussions

(0)

No comments yet

Be the first to share your thoughts!