Asymptotic Analysis of Empirical Dynamic Programming in Infinite-Horizon Stochastic Optimal Control
arXiv:2607.21520
2026
Theory
1 ideas extracted · analyzed Aug 30, 2026
What the math gives to ML
The paper provides a principled way to propagate finite-sample Bellman-operator error through an infinite-horizon fixed point, rather than treating value-target noise as independent across iterations. Its transferable asset is the decomposition into a static empirical-process perturbation and a discounted closed-loop linear response, which yields a state-correlated Gaussian approximation under a unique optimal policy. In neural fitted value iteration or model-based reinforcement learning, this can become an uncertainty estimator and a variance-aware target weighting scheme: estimate one-step Bellman uncertainty from bootstrap noise, then amplify it through the discounted transition resolvent. The nonunique-policy result also motivates a practical policy-instability diagnostic, since Gaussian uncertainty can fail near action ties.
Ideas from this paper
✗ Failed on benchmark
2026
Attach uncertainty to neural value targets by estimating the empirical one-step Bellman perturbation and propagating it through the discounted closed-loop transition operator. Use the resulting uncertainty to downweight high-variance Bellman targets or regularize the critic toward conservative predictions, especially in offline or model-based reinforcement learning.
Useful7/10
Difficulty6/10
Novelty6/10