Self-fictitious-play for Potential Monotone Ergodic Mean-field Games

arXiv:2608.15258 2026 Sampling 1 ideas extracted · analyzed Sep 1, 2026

What the math gives to ML

The paper supplies a constructive self-interacting stochastic process: an agent repeatedly computes a best response to a slowly changing belief, while that belief is updated from the agent's own occupation measure. The transferable asset is the separation of time scales together with a quantitative error law: the invariant behavior is within Wasserstein distance of the target equilibrium with error proportional to the square root of the belief-update rate. A practical neural-network adaptation is an adaptive latent-space sampler or replay distribution whose controller is a learned policy and whose occupancy belief is updated from the controller's own trajectories, rather than from an externally fixed proposal. This is most promising when mode coverage or distributional diversity matters and can be tested against ordinary Langevin sampling or a fixed replay mixture.

Ideas from this paper

Unverified 2026

Self-occupancy fictitious-play sampler

Turn the paper's self-fictitious-play process into a learned sampler for latent training examples or diffusion states. A controller network generates trajectories using a best response to a slowly updated occupancy belief, and the belief is updated from the controller's own states with an exponential occupation-measure update. The slow update prevents abrupt feedback loops while the controller continually adapts toward underrepresented regions.

Useful6/10
Difficulty6/10
Novelty7/10
Paper: Self-fictitious-play for Potential Monotone Ergodic Mean-field Games arXiv:2608.15258