Asymptotic Analysis of Empirical Dynamic Programming in Infinite-Horizon Stochastic Optimal Control

arXiv:2607.21520 2026 Theory 1 ideas extracted · analyzed Aug 30, 2026

What the math gives to ML

The paper provides a principled way to propagate finite-sample Bellman-operator error through an infinite-horizon fixed point, rather than treating value-target noise as independent across iterations. Its transferable asset is the decomposition into a static empirical-process perturbation and a discounted closed-loop linear response, which yields a state-correlated Gaussian approximation under a unique optimal policy. In neural fitted value iteration or model-based reinforcement learning, this can become an uncertainty estimator and a variance-aware target weighting scheme: estimate one-step Bellman uncertainty from bootstrap noise, then amplify it through the discounted transition resolvent. The nonunique-policy result also motivates a practical policy-instability diagnostic, since Gaussian uncertainty can fail near action ties.

Ideas from this paper

Failed on benchmark 2026

Bellman-Resolvent Uncertainty Targets

Attach uncertainty to neural value targets by estimating the empirical one-step Bellman perturbation and propagating it through the discounted closed-loop transition operator. Use the resulting uncertainty to downweight high-variance Bellman targets or regularize the critic toward conservative predictions, especially in offline or model-based reinforcement learning.

Useful7/10
Difficulty6/10
Novelty6/10
Paper: Asymptotic Analysis of Empirical Dynamic Programming in Infinite-Horizon Stochastic Optimal Control arXiv:2607.21520