Papers/2609.19170
🧪 Test?View on arXiv

Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes

Not provided in the abstract

temporal-difference learningreinforcement learningstability analysis
2609.19170
Builder Relevance
70%
2h ago

Abstract

This paper introduces regularized emphatic temporal-difference learning (RETD) to achieve stability under constant stepsizes in off-policy TD learning.

Reality Card

Core Claim

RETD achieves almost-sure convergence with harmonic diminishing stepsizes and demonstrates a conditional constant-stepsize moment-contraction result.

Method / Result

RETD has certified negative exponents on a two-state construction and one Baird point.

Limitations

The positive Baird ETD sign remains numerical, which may affect reproducibility.

Paper to code

Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.

No verified implementation link has been attached yet. AIBuzzHub will keep this panel separate from unverified search results.
← Back to all papers