Supervisory Control with Event Forcing Under Partial Observation
arXiv:2607.21040
2026
Dynamics
2 ideas extracted · analyzed Aug 30, 2026
What the math gives to ML
The paper provides a constructive partial-observation mechanism: a supervisor must choose identical forcing decisions for all strings that the supervisor cannot distinguish, while those decisions must preempt every specification-violating transition. Its key transferable asset is a feasibility test stronger than ordinary forcibility: for each observation class, the admissible forcing decisions must have a nonempty intersection across all compatible plant states. This transfers naturally to safety shields for partially observed neural controllers, where a learned policy proposes an intervention but a belief-state monitor must select one action or forcing mask that is safe for every latent state consistent with the observation. The non-closure under union also warns against independently combining locally safe neural intervention policies without rechecking joint consistency.
Ideas from this paper
△ Mechanism confirmed, baseline not beaten
2026
Add a discrete-event safety shield between a partially observed neural policy and the environment. The policy proposes a forcing action, but the shield permits it only when the same decision is safe for every latent plant state compatible with the current observation; otherwise it returns a certified inconsistency or a conservative fallback. This converts forcing consistency into an implementable robust action-selection rule rather than trusting a single estimated hidden state.
Useful8/10
Difficulty6/10
Novelty7/10
△ Mechanism confirmed, baseline not beaten
2026
Train a recurrent policy or neural controller so that histories with the same observation are forced toward the same intervention decision, while simultaneously requiring that the shared decision covers all unsafe latent transitions. This is stronger than ordinary action imitation or latent-state consistency because the loss explicitly penalizes cases where two observationally indistinguishable histories demand incompatible safety actions.
Useful7/10
Difficulty5/10
Novelty7/10