Unverified
2026
Replace a dense unconstrained channel-mixing matrix with a differentiable product of exponentials of a few skew-symmetric generators and their iterated commutators. The resulting layer is exactly orthogonal, preserves feature norms, and can express rotations in directions not explicitly stored as independent parameters. This is especially suitable for residual MLP blocks, recurrent state transitions, and networks processing rotation- or pose-valued features.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Augment a neural field or neural operator with a bank of localized, divergence-free moving packets whose radius and amplitude follow the Hill scaling rather than ordinary Gaussian scaling. The packet coefficients can represent unresolved flow corrections while keeping their L^2 contribution approximately invariant under refinement, preventing fine-scale features from becoming numerically negligible or explosively large.
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Replace feature-only graph pooling with a relaxed spectral-minimal partition layer. The layer assigns nodes to k clusters while favoring clusters with large algebraic connectivity, producing coarsened nodes that are internally well connected and less likely to contain bottlenecks. The resulting pooled graph can be used by a hierarchical GNN or graph transformer.
Useful6/10
Difficulty5/10
Novelty5/10
Unverified
2026
Replace an MoE router's single softmax distribution with a normalized coordinate-wise product of several simplex-valued routing factors. The product preserves positivity and normalization but, as depth grows, concentrates mass on a small subset of experts, creating a mathematically controlled heavy-tailed routing prior rather than relying only on an auxiliary load-balancing loss.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Replace an unconstrained spatial residual block by a discretized transport evolution whose generator is skew-adjoint. Symmetric channel matrices and divergence-free spatial coefficients make the continuous operator energy-preserving, while a matrix exponential or Cayley transform gives an exactly norm-preserving discrete layer.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Add a stable linear latent state-space block whose controllability Gramian is trained toward a chosen positive-definite target using squared Bures–Wasserstein distance. Direction-specific semidefinite constraints can suppress disturbance amplification in nuisance coordinates while preserving controllability in coordinates needed for prediction.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace or augment the coordinate embedding of a neural operator, PINN, or coordinate MLP with Chebyshev features plus rational features whose poles are selected by the AAA rational approximation algorithm. The rational features should represent boundary layers and other localized singular structures with fewer channels than a high-degree polynomial basis, reducing Gibbs-like oscillations and improving accuracy at small diffusion-to-advection ratios.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Replace an unconstrained multiscale residual block by the sum of a fractional diffusion branch and a drift or transport branch whose strength follows the PDE scaling law. At finer spatial scales, the drift coefficient is multiplied by R^{2s-1}; this suppresses unstable transport when s>1/2 while preserving equal-strength diffusion and drift at the critical value s=1/2.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Replace the generic nonlinear drift in a two-dimensional continuous-time recurrent cell by a learnable piecewise-linear Lienard restoring force. Fold breakpoints and jump breakpoints become explicit architectural controls for creating multiple oscillatory attractors, allowing hidden states to encode phase, mode, or periodic memory. Weak input coupling can select or perturb attractors while preserving the autonomous cycle structure.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace an unconstrained dense transition or recurrent matrix with a normal matrix $A=U\operatorname{diag}(\lambda)U^*$, where $U$ is unitary and $\lambda$ contains learnable eigenvalues. The layer can be initialized by fitting a normal matrix to input-output pairs through the paper's objective, then trained with Riemannian updates that keep $U$ unitary and preserve the normal-operator structure.
Useful6/10
Difficulty6/10
Novelty6/10
Unverified
2026
Track the variance of information content in a neural representation or routing distribution and use its interior maximum as a data-driven transition signal. The monitor distinguishes collapse, where nearly all probability occupies one state, from unstructured noise, where all states are equiprobable; both have low complexity, while structured intermediate distributions have high complexity.
Useful6/10
Difficulty4/10
Novelty7/10
Unverified
2026
Replace an unconstrained recurrent transition with block-diagonal planar rotations whose angles are learned or conditioned on a slowly varying context variable. The resulting hidden-state norm and each two-dimensional block energy are exactly invariant in the ideal recurrence, preventing exploding or vanishing recurrent dynamics while retaining phase information over long horizons.
Useful6/10
Difficulty4/10
Novelty4/10
Unverified
2026
Add a Renyi divergence penalty between the current network output distribution and a frozen reference distribution representing the pretrained model, a teacher, or a retained-data equilibrium. The Renyi order k becomes a control parameter: k greater than 1 strongly penalizes examples on which the new model assigns disproportionately more probability than the reference, while orders below 1 emphasize support mismatch and low-probability regions. Sweep or anneal k and detect a transition between…
Useful6/10
Difficulty4/10
Novelty6/10
Unverified
2026
Construct a recurrent or state-space layer whose equilibrium Jacobian is placed near a nondegenerate Bogdanov–Takens point, then use a small unfolding parameter to move between damped, oscillatory, and slowly relaxing regimes. Unlike eigenvalue-only initialization near one, this controls both the double-zero center structure and the quadratic nonlinear coefficients that determine the local phase portrait.
Useful6/10
Difficulty6/10
Novelty8/10
Unverified
2026
Partition a long integration interval into M short segments and assign one neural trajectory approximator to each segment. Instead of asking a single network to satisfy the ODE and initial condition over the entire horizon, construct every segment so that its value at the left boundary is exactly the terminal value predicted by the previous segment. This removes interface discontinuities from the optimization problem and should improve long-horizon trajectory accuracy and gradient stability.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Constrain a low-dimensional neural state-space model so that its vector-field Hessian approximately satisfies the paper's Pascal-Hessian condition. Combine the resulting latent dynamics with an observer correction driven by the prediction residual, giving a model whose hidden-state estimation error can be assigned a desired linear decay rate.
Useful6/10
Difficulty7/10
Novelty8/10
Unverified
2026
Add a Jacobian cone-field regularizer to recurrent dynamics so that tangent directions expand and remain aligned with an unstable cone outside a designated critical neighborhood. The network is not forced to be uniformly expanding: the regularizer is disabled near the critical set, allowing controlled bifurcation-like behavior while exposing where long-horizon sensitivity changes.
Useful6/10
Difficulty7/10
Novelty8/10
Unverified
2026
For a neural block with matrix-valued activations and transformation Y = A X B, regularize the exact coupled spectrum of the two-sided map instead of penalizing A and B independently. A large singular direction in A is penalized more strongly when the corresponding singular direction in B is also large, directly controlling joint feature amplification.
Useful6/10
Difficulty5/10
Novelty5/10
Unverified
2026
Add a stochastic pairwise regularizer that penalizes only coordinate pairs whose normalized neural-field difference exceeds a threshold. Unlike a conventional fractional Sobolev penalty, the weak-type functional uses an indicator and a distance weight, and its Gamma-limit guarantees convergence toward a local gradient energy as the threshold grows.
Useful6/10
Difficulty4/10
Novelty7/10
Unverified
2026
Attach each token or graph node a learned scalar charge q_i and add a fractional stable kernel K_ij = exp(-tau D |q_i-q_j|^alpha) to the interaction mechanism. Constrain 0 < alpha <= 2, the exact range in which the kernel is positive semidefinite for arbitrary finite real charge sets, and optionally make tau layer-dependent to obtain multiscale interactions. This provides a principled alternative to unconstrained learned distance biases and can be used either as an attention-logit bias or as a…
Useful6/10
Difficulty4/10
Novelty5/10
Unverified
2026
Construct a neural mixing layer only from Brauer generators for the orthogonal group: identity, pairwise swaps, and pairwise contractions with the Euclidean metric. This gives an exactly O(2)-equivariant alternative to unconstrained tensor mixing, with trainable coefficients but fixed symmetry-preserving basis maps.
Useful6/10
Difficulty5/10
Novelty5/10
Unverified
2026
Construct training and evaluation examples whose every (d-1)-variable marginal is exactly independent, but whose full d-variable distribution contains a parity interaction. This isolates genuine high-order reasoning from shortcuts based on pairwise or lower-order statistics and can expose whether attention or MLP architectures learn the intended interaction.
Useful6/10
Difficulty3/10
Novelty6/10
Unverified
2026
Build each nonlinear correction in an inverse neural operator from explicit bilinear products of learned operator features, following the inverse Born expansion instead of using an unconstrained pointwise MLP. Use a square activation to implement multiplication exactly, and truncate the interaction order so the model has a controllable polynomial structure.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Add a hybrid geometric penalty between an attention distribution at one layer or training step and a reference distribution, such as detached attention from the preceding layer or optimization step. The penalty allows attention mass to move between nearby token positions at a transport cost while separately charging for local compositional changes, producing a structured alternative to KL or entropy regularization.
Useful6/10
Difficulty7/10
Novelty7/10