Failed on benchmark
2026
Poisson-Calibrated Candidate-Pool Scheduler
Replace a fixed number of randomly sampled continuous actions with a state-dependent candidate pool whose size is chosen from the predicted extreme-value error of the best candidate. If the local action deficit has order \(\|u-u^\star\|^\kappa\) in an effective dimension \(d\), the best sampled action has expected Bellman error proportional to \(N^{-\kappa/d}\). This gives an explicit stopping rule for increasing the pool only when the estimated residual action error is larger than the…
Paper: Poisson Tangent Limits and Critical Policy Switching for Sampled Bellman Operators
arXiv:2608.11549