NobleBlocks
Public

Regularized Gradient Temporal-Difference Learning

Published in arXiv (Cornell University) • Jan 28, 2026
NobleIDNI0P25W09R19S81
Authors:
Hyunjun Na
,
Donghwan Lee

Abstract

Gradient temporal-difference (GTD) learning algorithms are widely used for off-policy policy evaluation with function approximation. However, existing convergence analyses rely on the restrictive assumption that the so-called feature interaction matrix (FIM) is nonsingular. In practice, the FIM can ...

Finding related papers...

Discussions

(0)

No comments yet

Be the first to share your thoughts!