Trajectory-Regularized Stochastic Optimal Control via KL Divergence

arXiv:2607.22201 2026 Regularization 1 ideas extracted · analyzed Aug 30, 2026

What the math gives to ML

The paper provides a constructive way to regularize a controlled stochastic trajectory toward a reference process: Girsanov's theorem converts trajectory-level KL divergence into an integral quadratic penalty on drift mismatch. This preserves dynamic programming and changes the local control curvature from the original cost matrix to an augmented matrix containing the reference-dynamics metric. The most direct neural-network transfer is a reference-preserving policy or world-model objective in which deviations are penalized in noise-whitened drift space rather than by an uncalibrated action-space norm, with the regularization strength controlled by a measurable trajectory-KL budget.

Ideas from this paper

Failed on benchmark 2026

Noise-Whitened Trajectory-KL Policy Regularization

Train a neural policy against task cost while penalizing its induced drift mismatch from a reference policy or offline-data dynamics model. Unlike action-space behavior cloning, the penalty weights deviations by the inverse diffusion covariance, so deviations in highly noisy directions are cheap and deviations in predictable directions are expensive. This gives a principled interpolation between reference preservation and task optimization.

Useful7/10
Difficulty5/10
Novelty6/10
Paper: Trajectory-Regularized Stochastic Optimal Control via KL Divergence arXiv:2607.22201