Asymptotic Analysis of Empirical Dynamic Programming in Infinite-Horizon Stochastic Optimal Control
Abstract
We derive statistical limit theorems for sample-based approximations of infinite-horizon discounted stochastic optimal control problems in discrete time. Our first result is a functional central limit theorem for the sample-based value function under a uniqueness-type condition on population optimal policies. The limiting law is a mean-zero Gaussian process characterized by a linear fixed point equation that resembles a dynamic programming principle. We compare these asymptotics with those obtained from sample-based policy optimization and illustrate that their limiting variances can be different. We also derive a limit theorem for models with nonunique optimal policies, where the limiting law may be non-Gaussian. Applications to inventory control and renewable harvesting illustrate the theory.
Disclosure
“xample using subsampling methods along the lines of those studied for two-stage stochastic programming in [6]. Acknowledgments We are grateful to Prof. Alexander Shapiro for many insightful discussions. The authors acknowledge the use of ChatGPT 5.5 and 5.6, and Claude (Opus 4.8) for assistance with language editing, organization, and presentation. The authors reviewed and revised all generated text and take responsibility for the final content. A. Proofs for Section 2 A.1. Proo”
PDF page 21
- Classification
- Rewriting existing author-written text
- Multiplier
- 4
- Verified
Structural counts
Count notes
- Source counts use the expanded primary TeX file CLT-Infinite-Horizon.tex.
- Appendix pages include the first PDF page with an explicit Appendix heading through the final page.