Online Control via Counterfactual Tracking

arXiv:2607.13029 2026 Dynamics 1 ideas extracted · analyzed Aug 30, 2026

What the math gives to ML

The paper separates policy selection from physical execution: arbitrary causal policies are simulated counterfactually, while a fixed stabilizing feedback controller tracks their predicted state-input trajectory on the real plant. This provides a safety and robustness layer for an ensemble of neural policies without requiring shared architectures, memory lengths, or controller orders. The strongest neural-network adaptation is an online Hedge or PAC-Bayes selector over heterogeneous policies combined with a linear error-feedback tracker. It is most applicable to online reinforcement learning or model-based control with a known or locally learned linearized plant.

Ideas from this paper

Mechanism confirmed, baseline not beaten 2026

Counterfactual-tracking policy ensemble

Maintain a posterior over heterogeneous neural policies, simulate each policy on the same revealed disturbance sequence, and track a posterior-weighted counterfactual reference instead of directly switching among deployed policies. A stabilizing feedback correction keeps the physical state close to the reference, while exponential-weights updates favor policies with low counterfactual cost.

Useful7/10
Difficulty5/10
Novelty7/10
Paper: Online Control via Counterfactual Tracking arXiv:2607.13029