Optimal Constrained sc-LTL Planning in MDPs via Switching Policies

arXiv:2608.05021 2026 Architecture 1 ideas extracted · analyzed Aug 31, 2026

What the math gives to ML

The paper provides a constructive method for converting non-Markovian temporal specifications into Markovian control by augmenting each environment state with deterministic automaton states. Its most transferable mechanism is an irreversible switching policy: use a policy for unresolved objectives, then switch to a policy specialized for the remaining objective after one target is reached. This can become a neural constrained-RL architecture with exact symbolic memory and separate policy heads, reducing interference between behaviors before and after a goal or safety event. The clearest test is sparse-reward navigation with reachability and safety constraints.

Ideas from this paper

Unverified 2026

Automaton-Gated Objective Switching

Augment a neural policy with deterministic DFA states for the task objective and safety constraint, then select among objective-specific policy heads using those states. Before either target is reached, execute a mixed policy; after one target is reached, switch permanently to the policy specialized for the remaining target.

Useful6/10
Difficulty4/10
Novelty5/10
Paper: Optimal Constrained sc-LTL Planning in MDPs via Switching Policies arXiv:2608.05021