Data-Driven optimal control via Koopman operators and Hamilton-Jacobi-Bellman equations

arXiv:2608.11808 2026 Dynamics 1 ideas extracted · analyzed Sep 1, 2026

What the math gives to ML

The paper combines Koopman-generator identification with geometric solution of Hamilton-Jacobi-Bellman equations, offering a structured way to learn controlled dynamics and value functions from trajectories. The transferable asset is the generator identity: derivatives of observables along trajectories become a differentiable, data-estimable representation of the vector field and control directions. A practical neural adaptation is to learn a control-affine latent lift and train a value network with an HJB residual computed through that learned generator, producing feedback directly from the value gradient. This may improve data efficiency and closed-loop stability relative to an unconstrained actor-critic, especially in low-dimensional continuous control.

Ideas from this paper

Mechanism confirmed, baseline not beaten 2026

Koopman-generator HJB critic

Replace an unconstrained learned dynamics model in model-based reinforcement learning or neural optimal control with a Koopman-style observable lift and an explicitly estimated infinitesimal generator. Train a value network against an HJB residual formed from this generator, so the critic is constrained by the observed vector field and control directions rather than relying only on temporal-difference targets.

Useful7/10
Difficulty5/10
Novelty6/10
Paper: Data-Driven optimal control via Koopman operators and Hamilton-Jacobi-Bellman equations arXiv:2608.11808