✗ Mechanism failed
2026
Represent Q-values using latent coefficients and a convex reconstruction operator rather than an unconstrained linear head. Enforce that reconstruction and compression are sup-norm nonexpansive, so the approximate Bellman operator remains a gamma-contraction and cannot exhibit the usual linear-function-approximation divergence.
Useful7/10
Difficulty5/10
Novelty7/10
✗ Mechanism failed
2026
Regularize a recurrent or state-space model using finite-time Lyapunov exponents of its actual hidden-state transition products. Penalize collapsed adjacent exponents while also controlling the largest exponent, encouraging several useful state directions instead of one dominant direction or universal contraction.
Useful7/10
Difficulty6/10
Novelty6/10
✗ Mechanism failed
2026
Train on a sequence of Jin–Xin relaxation problems with decreasing relaxation width rather than training immediately on the singular conservation law. The network predicts both the conserved state and an auxiliary flux, and each stage is initialized from the previous stage so that the learned shock profile sharpens gradually.
Useful7/10
Difficulty5/10
Novelty7/10
✗ Mechanism failed
2026
Replace an explicit Euler residual update for a skew-coupled hidden state with a five-stage palindromic composition of exact shear maps. Use a=1/4, the unique real coefficient maximizing the analyzed spectral CFL interval, and adapt the step size from an estimate of the learned coupling operator's spectral norm.
Useful7/10
Difficulty5/10
Novelty6/10
✗ Mechanism failed
2026
Replace an opaque MLP vector field with a stack of trainable symbolic primitives that can express linear terms, monomials, products, and related analytic operations. Apply an L1 penalty and prune small primitive coefficients after rollout training, yielding a compact dynamics module that is cheaper to evaluate and easier to inspect.
Useful7/10
Difficulty6/10
Novelty6/10
✗ Mechanism failed
2026
Replace a directed sequence-memory chain with a circular recurrent state propagated by a learned delayed convolution. The same learned kernel can support forward and reverse replay because replay direction is a dynamical mode of the ring, rather than requiring plasticity to explicitly learn both forward and backward synapses.
Useful7/10
Difficulty5/10
Novelty8/10
Unverified
2026
Replace a local smoothness penalty or local state transition along a sequence or depth coordinate by a marginal fractional quadratic energy with Fourier multiplier |k|. The sigma=1 kernel is nonlocal and scale-free, so it can preserve long-range correlations while suppressing high-frequency instability more selectively than an ordinary Laplacian penalty.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Apply the paper's empirical preimage-entropy construction to a learned recurrent transition map, penalizing excessive distinguishable hidden-state histories that produce the same current state while preserving multiple histories when the task requires genuine multimodality. Unlike a raw inverse-Jacobian penalty, the regularizer is computed only among inverse trajectories having similar empirical state distributions, so it distinguishes useful multimodal memory from uncontrolled branch explosion.
Useful6/10
Difficulty7/10
Novelty9/10
Unverified
2026
Treat the empirical hidden-state distribution of a recurrent or state-space model as a Wasserstein-space state and estimate the linearized pushforward operator on perturbation vector fields. Penalize tangent modes whose estimated transfer gains exceed one, while retaining near-unit fixed modes that represent robust invariant distributional structure.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace an unconstrained residual block by a first-order gradient-flow correction whose energy contains first-, second-, and third-difference penalties, mirroring the paper's higher-gradient gravitational energy. The correction suppresses high-frequency modes while retaining a trainable nonlinear residual branch, and its step size can be chosen from an explicit spectral stability bound.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Add a pressure-like recurrent state to a neural surface-flow decoder and update it from the predicted local divergence, creating a learned or fixed feedback loop that drives vector outputs toward local incompressibility. Unlike a static divergence penalty, the state can accumulate constraint violations and produce corrective tangent gradients at each refinement step.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Apply an inverse-square Calogero barrier to the eigenvalues of a recurrent or state-space transition Jacobian, discouraging unstable eigenvalues and pathological eigenvalue collisions without forcing the matrix to be Hermitian. The paper's non-Hermitian scattering picture motivates treating the spectrum as correlated rather than assuming an ordinary pairwise Coulomb gas; the inverse-square term is used as a local, computable surrogate for that mechanism.
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Regularize the state-transition or input-output Jacobian of a recurrent, state-space, or implicit neural network so that its complex eigenvalue cloud belongs to a selected non-Hermitian symmetry class and has the corresponding unfolded pair statistics. Combine this statistical-shape constraint with an explicit spectral-abscissa or spectral-radius margin, preventing the network from obtaining good average singular values while remaining highly non-normal and transiently unstable.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Equip a neural tracker with an explicit discrete posterior over candidate latent states, or approximate that posterior with particles or an ensemble, and monitor both its spread and its distance from the target or delayed supervision signal. Under likelihood-temperature misspecification, use the paper's two failure modes as a controller: flatten an overconfident posterior that is localized at the wrong state, while increasing observation trust when the posterior is diffuse but evidence is…
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Insert a sparse heavy-tailed nonlocal mixing operator into a residual sequence, graph, or spatial network so that information can traverse distant positions without stacking many local layers. Use the critical tail exponent s = 1/2, whose truncated first moment grows logarithmically and predicts an effective propagation distance proportional to depth times log depth rather than merely depth.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Replace an unconstrained recurrent or residual linear transition with a matrix generated through the paper's twisted Cayley chart and exact exponential flow. The layer evolves a constrained operator analytically rather than learning arbitrary weights, while retaining trainable symmetric chart coordinates and a continuous time-scale parameter.
Useful6/10
Difficulty6/10
Novelty6/10
Unverified
2026
Partition a neural network into heterogeneous parameter blocks or maintain several worker replicas, and model each block's optimizer state as a constrained linearized dynamical agent. At every synchronization interval, jointly optimize a finite sequence of parameter updates and a feasible common terminal parameter target, while enforcing consensus through distributed primal-dual iterations. Unlike ordinary gradient descent toward a fixed or implicit target, the target is selected together with…
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace the momentum update in a gradient optimizer by inertial motion plus a gradient-difference term, which discretely approximates Hessian-driven damping. Choose the damping coefficient and step size using the paper's refined stability inequality instead of the older restrictive bound, and adapt them whenever the estimated smoothness changes.
Useful6/10
Difficulty4/10
Novelty5/10
Unverified
2026
Parameterize an orthogonal or semi-orthogonal neural weight matrix directly on the Stiefel manifold and update it with a Cayley retraction instead of unconstrained SGD plus a penalty or QR projection. The update preserves orthogonality exactly, is second-order accurate for the appropriate metric, and avoids the cubic QR factorization at every optimizer step.
Useful6/10
Difficulty5/10
Novelty5/10
Unverified
2026
Add a fixed, spatially correlated perturbation field to every layer of a CNN or 2D state-space model, with the perturbation decomposed into transverse and longitudinal Fourier components. Unlike ordinary injected noise, the same field is reused for all training examples and all forward passes, allowing it to act as a structured architectural flow that can promote global feature alignment. Sweep the transverse fraction at fixed total perturbation variance and test for the predicted ordering…
Useful6/10
Difficulty6/10
Novelty8/10
Unverified
2026
Replace a neural controller's pointwise action outputs over a finite horizon with Bernstein control points whose convex hull satisfies actuator and trajectory constraints. The network predicts the control points, while a robust margin accounts for bounded tracking or model-prediction error, making continuous-time actuator feasibility checkable from finitely many inequalities.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Construct a generative or recurrent neural architecture with several contractive or mildly expanding branches, and explicitly control the geometric complexity of its invariant set using the sub-additive singular-value pressure of branch-Jacobian products. Instead of regularizing only the operator norm, the model can preserve anisotropic directions while targeting a desired attractor dimension, potentially improving coverage of structured data without uncontrolled folding or collapse.
Useful6/10
Difficulty6/10
Novelty8/10
Unverified
2026
Replace an ordinary local convolution or token-mixing block by two nonnegative feature populations A and B that diffuse and drift along the spatial or token axis, with transport rates increasing quadratically with local population and with an optional directional bias. Add a local A plus B to empty reaction so mutually conflicting feature mass is removed rather than merely averaged. The module should produce adaptive competition, and its isolated relaxation should exhibit a measurable t raised…
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Replace the dense output of selected linear projections with a two-sided magnitude threshold that emits zero for small values but preserves signed large values. Learn one positive threshold per projection, or optionally one threshold per output channel, so the network discovers where sparse events can be removed while retaining outlier information.
Useful6/10
Difficulty4/10
Novelty5/10