Unverified
2026
Use a small set of known or trusted pairwise costs to estimate the effective entropic temperature of a Sinkhorn attention or mixture-of-experts routing layer directly from its observed transport plan. This provides a calibration controller that can detect over-concentrated routing and adjust epsilon without backpropagating through a costly temperature search.
Useful6/10
Difficulty4/10
Novelty8/10
Unverified
2026
Replace the fixed delay in a temporal layer with a distribution of physically structured delays induced by uncertain transport velocity. The layer aggregates features arriving at several travel times and can use the deterministic mean-velocity path during most training steps, periodically correcting it with stochastic samples.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Run an ensemble of noisy optimization trajectories and regard trajectories that return to the same loss basin as competing dynamical phases. Estimate a complex return generating function from their path costs; a near-zero of this function signals cancellation between trajectory families and predicts an abrupt change in basin occupancy. Use the signal to reduce learning rate or optimizer noise near a transition, or increase noise when one phase dominates too early.
Useful6/10
Difficulty6/10
Novelty8/10
Unverified
2026
Replace a standard recurrent or neural-CDE Euler transition with a second-order rough transition that receives both first-order increments of the input path and learned second-order branched increments. Unlike a geometric signature block, the second-order coefficients are independent learned maps rather than being forced to equal derivatives or shuffle-symmetric combinations of first-order vector fields, allowing the model to represent order-sensitive and non-geometric interactions in irregular…
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Replace a local smoothness penalty or local state transition along a sequence or depth coordinate by a marginal fractional quadratic energy with Fourier multiplier |k|. The sigma=1 kernel is nonlocal and scale-free, so it can preserve long-range correlations while suppressing high-frequency instability more selectively than an ordinary Laplacian penalty.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Add a multiscale Besov penalty to the output of a shallow ReLU^k network, targeting the smoothness threshold that the paper proves is sufficient for finite ridge-variation representation. This suppresses pathological high-frequency output while preserving low-frequency approximation, providing a principled alternative to ordinary parameter weight decay.
Useful6/10
Difficulty4/10
Novelty6/10
Unverified
2026
Apply the paper's empirical preimage-entropy construction to a learned recurrent transition map, penalizing excessive distinguishable hidden-state histories that produce the same current state while preserving multiple histories when the task requires genuine multimodality. Unlike a raw inverse-Jacobian penalty, the regularizer is computed only among inverse trajectories having similar empirical state distributions, so it distinguishes useful multimodal memory from uncontrolled branch explosion.
Useful6/10
Difficulty7/10
Novelty9/10
Unverified
2026
Replace an unrolled constrained inner optimization in a meta-learning or hyperparameter-learning system with a KKT-based single-level layer. Instead of imposing primal-dual complementarity exactly from the first iteration, solve a sequence of relaxed problems with decreasing complementarity tolerances, making early optimization smoother and reducing failures caused by degenerate active-set geometry.
Useful6/10
Difficulty6/10
Novelty4/10
Unverified
2026
Treat discrete training events such as gradient-norm spikes, curvature changes, rejected steps, or minibatch outliers as jump channels and apply an event-specific parameter update map. The optimizer should be evaluated using both progress and the information cost of selecting the feedback map, because feedback may reduce loss fluctuations or improve adaptation without changing the average update magnitude.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Treat the empirical hidden-state distribution of a recurrent or state-space model as a Wasserstein-space state and estimate the linearized pushforward operator on perturbation vector fields. Penalize tangent modes whose estimated transfer gains exceed one, while retaining near-unit fixed modes that represent robust invariant distributional structure.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace unconstrained token-to-expert routing with a nonnegative mixture of a small set of feasible routing or communication patterns. Enforce resource limits using the paper's one-sided positive-contribution bound, which gives a conservative certificate without enumerating all joint token activation scenarios. Expand the pattern set only when a separation procedure finds a routing pattern that improves the router objective while adding useful capacity information.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Add a recurrent associative matrix to each selected transformer layer so recent key-value relationships can be retrieved without retaining every past token or performing gradient updates. The matrix uses input-dependent retention and write gates, but retrieval is always performed from the pre-write state, preventing the current target from leaking into its own prediction. Frobenius-norm clipping makes the recurrent memory bounded and provides a direct stability control.
Useful6/10
Difficulty4/10
Novelty4/10
Unverified
2026
Replace an unconstrained residual block by a first-order gradient-flow correction whose energy contains first-, second-, and third-difference penalties, mirroring the paper's higher-gradient gravitational energy. The correction suppresses high-frequency modes while retaining a trainable nonlinear residual branch, and its step size can be chosen from an explicit spectral stability bound.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Add a pressure-like recurrent state to a neural surface-flow decoder and update it from the predicted local divergence, creating a learned or fixed feedback loop that drives vector outputs toward local incompressibility. Unlike a static divergence penalty, the state can accumulate constraint violations and produce corrective tangent gradients at each refinement step.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Apply an inverse-square Calogero barrier to the eigenvalues of a recurrent or state-space transition Jacobian, discouraging unstable eigenvalues and pathological eigenvalue collisions without forcing the matrix to be Hermitian. The paper's non-Hermitian scattering picture motivates treating the spectrum as correlated rather than assuming an ordinary pairwise Coulomb gas; the inverse-square term is used as a local, computable surrogate for that mechanism.
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Regularize the state-transition or input-output Jacobian of a recurrent, state-space, or implicit neural network so that its complex eigenvalue cloud belongs to a selected non-Hermitian symmetry class and has the corresponding unfolded pair statistics. Combine this statistical-shape constraint with an explicit spectral-abscissa or spectral-radius margin, preventing the network from obtaining good average singular values while remaining highly non-normal and transiently unstable.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Equip a neural tracker with an explicit discrete posterior over candidate latent states, or approximate that posterior with particles or an ensemble, and monitor both its spread and its distance from the target or delayed supervision signal. Under likelihood-temperature misspecification, use the paper's two failure modes as a controller: flatten an overconfident posterior that is localized at the wrong state, while increasing observation trust when the posterior is diffuse but evidence is…
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Replace an unconstrained hidden-to-hidden interaction in an MLP or transformer feed-forward block by two gauge-related branches. Split channels with an orthogonal involution Θ, constrain the learned interaction K to anticommute with Θ, and use opposite signs of K in paired branches. This creates a testable inductive bias in which the learned interaction only transfers information between the two channel subspaces.
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Replace pointwise high-order derivative residuals in an eigenvalue PINN by an assembled dynamic-stiffness residual \(\mathbf W(\omega)q_\theta\), where each element matrix is obtained from homogeneous PDE solutions. The network predicts nodal degrees of freedom or element boundary traces, while the exact frequency-domain operator enforces the physics without differentiating the network multiple times with respect to coordinates.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
When a neural model defines a product distribution over coordinates, compute the discrepancy to a target product distribution from per-coordinate divergences using the exact tensorization rule instead of sampling full vectors and estimating a joint divergence. Implement KL, chi-squared, and squared-Hellinger variants as drop-in losses, with an optional learned choice among these mathematically tensorizable families.
Useful6/10
Difficulty3/10
Novelty6/10
Unverified
2026
Replace an unconstrained recurrent or residual linear transition with a matrix generated through the paper's twisted Cayley chart and exact exponential flow. The layer evolves a constrained operator analytically rather than learning arbitrary weights, while retaining trainable symmetric chart coordinates and a continuous time-scale parameter.
Useful6/10
Difficulty6/10
Novelty6/10
Unverified
2026
Partition a neural network into heterogeneous parameter blocks or maintain several worker replicas, and model each block's optimizer state as a constrained linearized dynamical agent. At every synchronization interval, jointly optimize a finite sequence of parameter updates and a feasible common terminal parameter target, while enforcing consensus through distributed primal-dual iterations. Unlike ordinary gradient descent toward a fixed or implicit target, the target is selected together with…
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace the noisy Hutchinson estimate of a neural-network Hessian trace with a variance-reduced Hutch++ estimate computed only from Hessian-vector products. Use the estimated normalized curvature to cap or rescale the optimizer step, so learning-rate reductions occur when the loss landscape becomes globally sharp rather than when an individual minibatch gradient happens to be large.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Replace the momentum update in a gradient optimizer by inertial motion plus a gradient-difference term, which discretely approximates Hessian-driven damping. Choose the damping coefficient and step size using the paper's refined stability inequality instead of the older restrictive bound, and adapt them whenever the estimated smoothness changes.
Useful6/10
Difficulty4/10
Novelty5/10