Geometrically Attracting Random Recurrent Layer
Source paper: On the Existence of Geometrically Attracting Measures for Iterated Function Systems with Varying Sets of Transformations arXiv:2608.29022 ⓘ · analyzed Sep 1, 2026
AI-generated research hypothesis, automatically tested. Not peer-reviewed.
Idea description
Replace a recurrent update by a time-inhomogeneous random choice among candidate maps, and regulate the candidate Jacobian gains so that the expected product of gains contracts geometrically. This should make hidden-state distributions forget their initial state even when the map family and selection probabilities vary over time, improving long-horizon stability without requiring every individual candidate map to be strongly contractive.
Formulas
Mathematical statement
Let Xₜ₊₁ = Fₜ,ᴵₜ(Xₜ), where Fₜ,ᵢ is candidate map i at time t and Iₜ is sampled with probability pₜ,ᵢ. Let Lₜ,ᵢ satisfy ||Fₜ,ᵢ(x) − Fₜ,ᵢ(y)|| ≤ Lₜ,ᵢ ||x − y||. A sufficient geometric-attraction condition is E[Πₜ₌₀ⁿ⁻¹ Lₜ,ᴵₜ] ≤ Cρⁿ for constants C < infinity and 0 < ρ < 1. For independent selections this reduces to Πₜ₌₀ⁿ⁻¹ Σᵢ pₜ,ᵢLₜ,ᵢ ≤ Cρⁿ. If μₙ and νₙ are the state distributions generated from two initial distributions, their bounded-Lipschitz distance should decay geometrically: d_BL(μₙ,νₙ) ≤ Cρⁿd_BL(μ₀,ν₀). In the neural-network implementation, Lₜ,ᵢ is estimated by the spectral norm of the candidate hidden-state Jacobian, and the logarithmic average gain is regularized below a target log contraction rate.
Implementation notes
Integrate the mechanism into an RNN or state-space block with K candidate maps, such as Fₜ,ᵢ(hₜ,xₜ) = tanh(Wₜ,ᵢhₜ + Uₜ,ᵢxₜ + bₜ,ᵢ), and a time-dependent categorical gate pₜ,ᵢ. At each training step, estimate each candidate gain L̂ₜ,ᵢ using one to three power-iteration steps on the hidden-to-hidden Jacobian. A cheaper upper bound is ||Wₜ,ᵢ||₂ because tanh is 1-Lipschitz. Pseudocode: compute pₜ; sample a candidate or use the probability-weighted mixture; update hₜ₊₁ = Fₜ,ᴵₜ(hₜ,xₜ); estimate all L̂ₜ,ᵢ; compute gₜ = log(Σᵢ pₜ,ᵢL̂ₜ,ᵢ + epsilon); add λ times the squared positive part of gₜ − log(ρ*) to the task loss. Track γ̂ₜ = T⁻¹Σₜgₜ, and optionally rescale recurrent matrices when γ̂ₜ exceeds zero. The paper supplies the contraction mechanism; Jacobian estimates, routing probabilities, and finite-horizon exponents are measured empirically. First test on sequential MNIST and the adding problem against a GRU and an unconstrained random-map RNN with matched parameter counts. Run two identical input sequences from hidden states separated by unit norm, and record dₜ = ||hₜ − h′ₜ|| for 200 steps. The prediction is a stability boundary near γ̂ = 0: below it, log dₜ has negative slope approximately γ̂; above it, distances do not decay geometrically. The measured slope should agree with the accumulated log-gain within about 20 percent.
Verification
This idea has not been verified yet.
Verification happens in two stages: Stage 1 — a mechanism check on a toy system confirms the claimed mathematical phenomenon reproduces; Stage 2 — a benchmark implements the idea on a real (small) neural network task and compares it against a tuned baseline over 8 paired seeds with a permutation test.
Artifacts
Artifacts unavailable.