Unverified
2026
Distill a large or accurate latent transition model into a smaller discrete-state recurrent model while penalizing both its one-step transition mismatch and its lack of contraction. The filtering perturbation bound predicts that reducing the Dobrushin coefficient prevents errors from accumulating over long sequences, while reducing the transition discrepancy lowers the irreducible steady-state error.
Useful6/10
Difficulty4/10
Novelty7/10
Unverified
2026
Construct a mixture-of-experts layer whose experts compete for a normalized routing resource, and regularize the router so that every expert can grow when introduced at low abundance into the equilibrium dominated by any other expert. The ecological mutual-invasibility criterion becomes a quantitative anti-collapse condition: if expert B has positive invasion growth against expert A's equilibrium and A has positive invasion growth against B, neither single-expert state is locally stable against…
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace an unconstrained repeated averaging or message-passing operator by an average of positive isometric group actions whose mixing distribution satisfies the paper's bounded angular ratio condition. The resulting operator is Ritt, giving a mathematically certified bound on successive iterates and convergence of repeated application. This can stabilize deep equivariant stacks and reduce oscillatory feature dynamics.
Useful6/10
Difficulty4/10
Novelty7/10
Unverified
2026
Add a distribution-level regularizer that compares augmented second-moment matrices of neural activations using the affine-invariant Riemannian metric on SPD matrices. This aligns means, variances, and selected nonlinear moments while remaining invariant to invertible linear reparameterizations of feature coordinates.
Useful6/10
Difficulty5/10
Novelty5/10
Unverified
2026
Parameterize a cell-complex neural network by features on p-cells and derive lower-dimensional boundary features using the cellular boundary map over F2. For a 2D square complex, neighboring plaquette bits determine each link feature through XOR, reproducing the paper's exact gauge-law reconstruction and preventing the network from representing inconsistent open boundary configurations.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Construct intermediate training examples along an optimal-transport coupling between two strongly log-concave endpoint distributions, and regularize the network so that its output variance on each intermediate distribution is no larger than the sharp endpoint-interpolated Poincare scale times its expected input-Jacobian energy. This converts the paper's distributional inequality into a path-wise smoothness constraint for logits, embeddings, or scalar losses.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Use multiple independently initialized training replicas to detect discontinuous transitions in the learned state as a hyperparameter changes. A saddle-node event is identified when two locally stable or unstable solution branches collide, producing an abrupt jump in a validation-relevant order parameter; pseudo-arclength continuation can map this event and choose a hyperparameter path that avoids catastrophic branch loss.
Useful6/10
Difficulty6/10
Novelty8/10
Unverified
2026
Add a functional-calculus regularizer to the transition operator of an RNN, linear state-space model, or deep-equilibrium layer. The regularizer uses polynomial probes to detect non-normal transient amplification that ordinary eigenvalue-radius penalties can miss.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Add a differentiable regularizer to neural networks that learn sparse Fourier coefficients or trainable Fourier-feature frequencies. It penalizes predicted energy just outside the training interval when that energy exceeds the theorem-shaped envelope relative to observed in-domain L2 energy, discouraging cancellation patterns that fit the observed interval but explode nearby.
Useful6/10
Difficulty3/10
Novelty8/10
Unverified
2026
Replace a single smooth inverse predictor near detected ambiguity boundaries with multiple prediction branches and a soft gate. The gate is trained to preserve distinct decompositions rather than forcing the network to interpolate through a thin high-curvature transition layer, while a Jacobian or curvature penalty identifies unresolved ambiguity regions.
Useful6/10
Difficulty6/10
Novelty6/10
Unverified
2026
Represent input or parameter uncertainty locally by a low-order polynomial expansion of the network output, and compute only task-relevant directional third- and fourth-order moments. Add a penalty that calibrates or controls projected skewness and kurtosis, allowing the model to represent bent or elongated confidence regions without constructing a full dense moment tensor.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Use the conservation-law density to weight diffusion training examples by noise level instead of relying on uniform, cosine, or manually selected SNR weighting. This emphasizes noise regions whose local information contribution is largest while clipping the weights to prevent rare regions from destabilizing optimization.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Use the paper's singular stopping-gain term to explicitly measure how much learned feature covariance crosses a max or routing boundary. Penalize excessive covariance in the normal direction to the switching surface, rather than pretending that the max operation has an ordinary Hessian.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Add a targeted barrier or hinge loss to an existing attention or graph-mixing matrix that penalizes violations of signed circular-minor inequalities. Instead of enforcing only generic entrywise positivity, constrain higher-order noncrossing interactions encoded by determinants. This can suppress pathological oscillatory mixing while still allowing individual entries to be negative when the global structured sign pattern permits them.
Useful6/10
Difficulty4/10
Novelty8/10
Unverified
2026
Use midpoint or running ergodic averages of adversarial iterates for evaluation and checkpointing instead of exposing a single phase-dependent iterate. The mathematical attenuation factor suppresses rotational error, especially for modes with large step-size-times-frequency product.
Useful6/10
Difficulty2/10
Novelty4/10
Unverified
2026
Model a finite training run as a driven stochastic process whose control parameter is the learning rate or another scheduled hyperparameter. Compare the distribution of parameter perturbations, activations, logits, or losses after a finite-rate update to a reference distribution generated by a much slower approximately adiabatic schedule; reduce the learning rate when the estimated relative entropy exceeds a calibrated threshold.
Useful6/10
Difficulty6/10
Novelty8/10
Unverified
2026
Replace a fixed ridge coefficient in a neural network's final head with a controller driven by inverse spectral mass and hard-edge mass. The head can remain weakly regularized when the feature spectrum is healthy, but automatically increases ridge strength when small eigenvalues signal a high-risk interpolation regime.
Useful6/10
Difficulty4/10
Novelty5/10
Unverified
2026
Replace a dense learnable Fourier multiplier with a low-parameter multiplier concentrated near the common zero set of two polynomial constraint symbols. A linear constraint together with a cubic constraint can produce straight or curved frequency loci, allowing the network to represent directional long-range structure while using far fewer spectral parameters than a full 3D frequency grid.
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Modify learning-rate or annealing schedules so that local improvement is not mistaken for convergence when different parameter blocks occupy incompatible global modes. Measure a local-consistency score and a global-coherence score separately; slow training whenever local consistency is high but global coherence remains low, allowing competing parameter domains to merge before cooling further.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Insert a differentiable intersection-body-inspired map on positive spherical feature fields. The map contracts high-order angular variation while leaving degree-two ellipsoidal structure neutral, providing a principled alternative to generic smoothing that does not erase global anisotropy.
Useful6/10
Difficulty6/10
Novelty6/10
Unverified
2026
Replace a fixed Fourier or spectral resolution in a neural operator or sequence model with a data-adaptive spectral cutoff. Keep only modes whose estimated signal energy exceeds the noise-amplification and discretization floor implied by the available number of trajectories and samples per trajectory. This should reduce overfitting to high-frequency sensor noise and preserve accuracy when the same model is deployed at different sampling resolutions.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Use the M phase-aligned parameterizations produced by cyclic reformulation as an empirical ensemble of neural dynamics rather than selecting one phase or averaging only predictions. Their centroid supplies a nominal model, while their convex hull defines a low-dimensional uncertainty set used for robust rollout training and uncertainty-aware inference.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Replace or augment entropy-based MoE load balancing with a structured concave utility over expert loads. The utility is the geometric mean of positive linear coverage factors, so it rewards underused directions strongly while exhibiting diminishing returns for already-covered directions. Positive coefficients can encode expert capacity, hardware placement, or expert groups.
Useful6/10
Difficulty3/10
Novelty7/10
Unverified
2026
Use sign choices over redundant gradient or adapter proposals to keep the accumulated residual update small in the coordinatewise maximum norm. Constrain the sign controller to preserve a positive projection onto the desired descent direction, so it suppresses coordinate spikes without completely canceling optimization progress.
Useful6/10
Difficulty6/10
Novelty7/10