Regularized Gradient Temporal-Difference Learning
Published in arXiv (Cornell University) • Jan 28, 2026
NobleIDNI0P25W09R19S81
Authors:,
Hyunjun Na
Donghwan Lee
Abstract
Gradient temporal-difference (GTD) learning algorithms are widely used for off-policy policy evaluation with function approximation. However, existing convergence analyses rely on the restrictive assumption that the so-called feature interaction matrix (FIM) is nonsingular. In practice, the FIM can ...
Finding related papers...
Discussions
(0)No comments yet
Be the first to share your thoughts!