Trajectory-Regularized Stochastic Optimal Control via KL Divergence
arXiv:2607.22201
2026
Regularization
1 ideas extracted · analyzed Aug 30, 2026
What the math gives to ML
The paper provides a constructive way to regularize a controlled stochastic trajectory toward a reference process: Girsanov's theorem converts trajectory-level KL divergence into an integral quadratic penalty on drift mismatch. This preserves dynamic programming and changes the local control curvature from the original cost matrix to an augmented matrix containing the reference-dynamics metric. The most direct neural-network transfer is a reference-preserving policy or world-model objective in which deviations are penalized in noise-whitened drift space rather than by an uncalibrated action-space norm, with the regularization strength controlled by a measurable trajectory-KL budget.
Ideas from this paper
✗ Failed on benchmark
2026
Train a neural policy against task cost while penalizing its induced drift mismatch from a reference policy or offline-data dynamics model. Unlike action-space behavior cloning, the penalty weights deviations by the inverse diffusion covariance, so deviations in highly noisy directions are cheap and deviations in predictable directions are expensive. This gives a principled interpolation between reference preservation and task optimization.
Useful7/10
Difficulty5/10
Novelty6/10