Unverified
2026
Replace independent additive noise on spatial feature maps with stochastic advection by divergence-free vector fields. The perturbation preserves spatial volume and feature mass, while the associated Stratonovich-to-Itô correction provides a tunable diffusion that preferentially damps high-frequency spatial fluctuations.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Place doubly stochastic stream mixing immediately before a quantizer, activation compressor, or latent bottleneck and jointly optimize task loss with estimated code length. The paper's entropy argument says that this linear mixing cannot increase differential entropy, so it can provide cross-stream representation capacity without an ideal entropy-rate penalty; the entropy bottleneck then learns which feature values deserve bits.
Useful6/10
Difficulty4/10
Novelty4/10
Unverified
2026
Replace an unconstrained deep residual recurrence by a discretized diffusion system over feature or token positions, with trainable source terms and analytically constrained boundary feedback. The state remains nonnegative under nonnegative inputs, while negative boundary gains enforce exponential decay of perturbations and prevent exploding activations in very deep stacks.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Use the OT spectral bound as a conditioning signal for optimizing parameters of a neural cost or inverse-OT objective. Adapt the parameter step size and add a covariance floor whenever the estimated Jacobian lower bound collapses, preventing optimization from entering regions where Sinkhorn outputs become insensitive to the learned cost.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Replace a dense translation-invariant interaction matrix with a positive-definite Toeplitz kernel K_n(e^f) whose log-spectrum is parameterized by a small number of Fourier coefficients with 1/|k| decay. Use the paper's explicit quadratic term as a spectral-volume budget, allowing long-range structure while discouraging uncontrolled determinant growth and ill-conditioning. Subtracting this term from a log-determinant regularizer leaves a residual intended to capture higher-order deviations from…
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
For models that predict probability distributions, replace the usual Wasserstein-2 loss or a Huber penalty on the final Wasserstein distance with a Huber penalty on quantile-by-quantile prediction errors. This suppresses gradients from localized outliers while retaining quadratic gradients on the majority of the distribution, which is useful for uncertainty prediction, histogram prediction, and distributional distillation.
Useful6/10
Difficulty3/10
Novelty5/10
Unverified
2026
Search sparse reservoir wiring in graph space rather than repeatedly testing every candidate with its full nonlinear dynamics. Use graph descriptors to predict validation accuracy and nonlinear feature selectivity, then spend exact simulations on candidates with high predicted performance or high surrogate uncertainty.
Useful6/10
Difficulty6/10
Novelty6/10
Unverified
2026
Replace ordinary absolute positional embeddings with coordinates on a learned flat torus and use dual-lattice Fourier characters as positional features. Control the covariance of the coordinate fundamental domain so that the paper's inequality guarantees a lower bound on the smallest nonzero positional frequency, preventing the learned periodic coordinate system from developing arbitrarily weak or nearly constant modes.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Replace an unconstrained recurrent latent transition with a map having one deliberately expanding angular coordinate and strongly contracting transverse coordinates. The construction should produce a bounded chaotic attractor with a reproducible stationary distribution while preventing uncontrolled expansion in the remaining hidden dimensions.
Useful6/10
Difficulty6/10
Novelty8/10
Unverified
2026
Replace unconstrained spectral mixing with a three-component triadic interaction whose strength is determined by the quadratic phase mismatch R(xi,xi_1). Near-resonant products receive high weight because their phases remain coherent, while strongly nonresonant products are attenuated. The resonance bandwidth can be fixed from the frequency grid or learned as a positive parameter.
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Add a sparse, interpretable fractional-dynamics layer to a neural world model: candidate terms are evaluated through weak projections, while both their support and continuous derivative orders are selected by validation error versus model complexity. This avoids forcing the model to choose from a dense fixed dictionary containing many nearly collinear fractional orders.
Useful6/10
Difficulty6/10
Novelty6/10
Unverified
2026
Constrain a recurrent or state-space transition matrix so that its eigenvalues avoid a configurable annulus around the unit circle. This creates a stable/unstable decomposition and should reduce the accumulation of numerical, quantization, and activation-update errors over long sequences while preserving controlled long-term memory.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Add a bounded routing state to an RNN, state-space model, or mixture-of-experts layer, with several neutral fixed points representing persistent modes. The state moves between modes when far from a fixed point but escapes each mode only polynomially when close to it, creating controllable long memory without setting a linear eigenvalue arbitrarily close to one. A temperature parameter selects between an entropy-rich phase using many modes and a low-entropy phase concentrated near one preferred…
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Add an asymmetric distillation loss that requires target or teacher observables to be approximable by source or student observables on only a 1-epsilon mass subset of a coupling. Unlike symmetric feature alignment, the student is penalized only for failing to reproduce target functions on well-matched mass, making the objective robust to outliers, label noise, and partial domain mismatch. The inner minimization allows each target observable to select its best source probe rather than forcing a…
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Replace random or k-means initialization of a k-expert router with a moment-based range finder on a calibration batch of hidden states. Estimate a low-dimensional second-moment subspace, enlarge it using one-free-index third-Hermite contractions, and fit the router's expert centroids and weights only in this resulting subspace. The router can then operate on projected hidden states while retaining an optional small residual adapter.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Represent a collection of neural directions as generators of a zonotope and reward the volume spanned by their subsets. The objective favors complementary, non-collapsed vectors rather than merely pairwise-separated vectors, making it suitable for attention heads, MoE expert signatures, or embedding prototypes. Use normalized generators and positive gates so the regularizer cannot be increased trivially by scaling.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Replace fixed graph message weights with a source-node activity gate that amplifies or suppresses every outgoing message from that node. Use the linearized epidemic growth condition to calibrate the residual propagation strength so that the dominant graph mode is near, but below, an explicitly chosen stability threshold rather than being determined accidentally by the graph spectrum.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Insert a proximal layer after a graph, mesh, or spherical convolution that groups all coordinates belonging to the same Laplacian eigenspace and applies one shared shrinkage gate to the whole group. Unlike coefficientwise spectral pruning, the result is unchanged if the eigenvectors inside a repeated eigenspace are rotated, preventing arbitrary basis-dependent feature selection.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Build neural computation graphs with explicitly phase-budgeted serial and parallel branches, treating serial compositions as SRG products and parallel residual branches as SRG sums. Allocate phase centers theta_i so that every loop or branch aggregate stays away from -1, enabling stability-aware architecture search and constructive control of branch gains.
Useful6/10
Difficulty7/10
Novelty8/10
Unverified
2026
Use unsquared hinge penalties when a neural objective must obey strict priority semantics, and treat squared hinges as approximate penalties rather than exact enforcement mechanisms. Add a residual monitor that detects when a finite weighted solve is still trading a higher-tier violation for lower-tier improvement, then switches to a sequential cascade or projection step.
Useful6/10
Difficulty4/10
Novelty7/10
Unverified
2026
Replace an unconstrained q-way polynomial or tensorized feature layer with separate decomposable and primitive interaction channels. The decomposable channel models interactions explainable as products of lower physical-weight feature blocks, while the primitive channel captures residual factors that cannot be represented by those products. This should reduce redundant high-order parameters and provide a controllable inductive bias for compositional or disentangled representations.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Run Metropolis-Hastings directly on neural architectures modulo permutations of structurally exchangeable hidden units, channels, or experts rather than treating every labelled representation as a distinct architecture. Correct the proposal ratio using representation-orbit sizes, so architectures with many internal symmetries receive the intended posterior mass.
Useful6/10
Difficulty6/10
Novelty5/10
Unverified
2026
Insert a differentiable Fourier-domain layer after a network predicts a symmetric strain field, projecting every frequency onto the subspace satisfying isotropic mechanical equilibrium. The projection is a closed-form least-squares correction, so the network cannot spend capacity representing large equilibrium violations and the resulting field is physically admissible by construction.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Replace a fixed confidence-threshold early-exit rule with a finite-horizon optimal-stopping policy over the model's evolving posterior confidence. The controller stops when the calibrated expected terminal error is no greater than the cost plus expected value of executing another neural block, permitting time-dependent and nonmonotone stopping regions.
Useful6/10
Difficulty5/10
Novelty5/10