Unverified
2026
Learn a branching hierarchy for tokens, examples, or experts by greedily relocating leaves to reduce T-Robinson violations. The resulting tree supplies hierarchical candidate sets for retrieval or MoE routing, allowing the model to search a small subtree instead of all items while adapting the hierarchy to learned representations.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Model dynamic routing as a multitype branching process: an active token of type d probabilistically creates child activations of type d'. Estimate the corresponding mean offspring operator and regulate its Perron root to a target reproduction rate, typically near one. This should make adaptive-depth or recursively routed networks use sparse computation without producing either rapidly vanishing paths or uncontrolled activation explosions.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Before or during graph-neural-network inference, use a tree dynamic program to choose a limited set of active computation nodes and upgraded message-passing edges. A node receives an embedding from an active landmark only when the selected path has enough residual communication radius, so the planner directly optimizes weighted coverage under a joint node-and-edge budget. The resulting active subgraph is then used by a sparse GNN or graph transformer.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Replace unrestricted global attention or purely local convolution by a sparse distance-dependent interaction graph on a two-dimensional feature map. The edge probability or attention prior decays as \(r^{-(2+\sigma)}\), and \(\sigma\) becomes an explicit architectural control knob: small \(\sigma\) supplies mean-field-like global mixing, intermediate \(\sigma\) supplies long-range Wilson–Fisher behavior, and \(\sigma>2\) approaches a short-range model. The architecture should be evaluated not…
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Approximate a dense symmetric interaction matrix in a neural layer by \(\widehat A=C\widehat M C^{\top}\), but compute the small core \(\widehat M\) from a two-sided sketched least-squares fit rather than from the landmark principal submatrix. This preserves signed or indefinite directions and avoids exploding outputs caused by an almost-singular \(A(I,I)\).
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Replace an unconstrained Fourier-domain linear mixer with a bank of positive spectral kernels and a max-times erosion aggregator. For a nonnegative Fourier magnitude f, each kernel produces a quotient response f/psi_k and the layer takes the pointwise supremum over kernels, giving exact positive homogeneity and monotonicity. This is most suitable as a drop-in spectral mixing block in a CNN, vision transformer, or state-space model.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Construct a neural layer as a sum of equivariant spectral operators at dyadic frequency scales, with each scale represented by a smooth learnable multiplier instead of an unconstrained dense spectral table. Enforce derivative and off-diagonal decay constraints so high-frequency components cannot create arbitrarily large or spatially nonlocal responses. On a discretized homogeneous space, this gives a multiresolution equivariant alternative to a generic graph filter or convolution kernel.
Useful6/10
Difficulty7/10
Novelty6/10
Unverified
2026
Replace a single plug-in top-k router decision with a confidence correspondence containing every router parameter candidate and sparse expert assignment that remains compatible with calibration and the current input. Project this set onto a hierarchy of expert groups and return the finest group-level decision supported by all surviving explanations; otherwise coarsen the route or abstain. Active endpoint bracketing evaluates only candidates that could still change the projected routing report.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Convert a neural operator block into a shared-weight iterative fixed-point refinement scheme that exploits repeated smoothing while avoiding repeated low-resolution projections. Compute all refinement steps at an overresolved latent bandwidth and apply the target-bandwidth projection only at the end, reducing the opportunity for unresolved frequencies to alias into retained channels.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace a dense unconstrained channel-mixing matrix with a differentiable product of exponentials of a few skew-symmetric generators and their iterated commutators. The resulting layer is exactly orthogonal, preserves feature norms, and can express rotations in directions not explicitly stored as independent parameters. This is especially suitable for residual MLP blocks, recurrent state transitions, and networks processing rotation- or pose-valued features.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Augment a neural field or neural operator with a bank of localized, divergence-free moving packets whose radius and amplitude follow the Hill scaling rather than ordinary Gaussian scaling. The packet coefficients can represent unresolved flow corrections while keeping their L^2 contribution approximately invariant under refinement, preventing fine-scale features from becoming numerically negligible or explosively large.
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Replace feature-only graph pooling with a relaxed spectral-minimal partition layer. The layer assigns nodes to k clusters while favoring clusters with large algebraic connectivity, producing coarsened nodes that are internally well connected and less likely to contain bottlenecks. The resulting pooled graph can be used by a hierarchical GNN or graph transformer.
Useful6/10
Difficulty5/10
Novelty5/10
Unverified
2026
Replace an MoE router's single softmax distribution with a normalized coordinate-wise product of several simplex-valued routing factors. The product preserves positivity and normalization but, as depth grows, concentrates mass on a small subset of experts, creating a mathematically controlled heavy-tailed routing prior rather than relying only on an auxiliary load-balancing loss.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Add a stable linear latent state-space block whose controllability Gramian is trained toward a chosen positive-definite target using squared Bures–Wasserstein distance. Direction-specific semidefinite constraints can suppress disturbance amplification in nuisance coordinates while preserving controllability in coordinates needed for prediction.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace or augment the coordinate embedding of a neural operator, PINN, or coordinate MLP with Chebyshev features plus rational features whose poles are selected by the AAA rational approximation algorithm. The rational features should represent boundary layers and other localized singular structures with fewer channels than a high-degree polynomial basis, reducing Gibbs-like oscillations and improving accuracy at small diffusion-to-advection ratios.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Replace an unconstrained multiscale residual block by the sum of a fractional diffusion branch and a drift or transport branch whose strength follows the PDE scaling law. At finer spatial scales, the drift coefficient is multiplied by R^{2s-1}; this suppresses unstable transport when s>1/2 while preserving equal-strength diffusion and drift at the critical value s=1/2.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Replace an unconstrained dense transition or recurrent matrix with a normal matrix $A=U\operatorname{diag}(\lambda)U^*$, where $U$ is unitary and $\lambda$ contains learnable eigenvalues. The layer can be initialized by fitting a normal matrix to input-output pairs through the paper's objective, then trained with Riemannian updates that keep $U$ unitary and preserve the normal-operator structure.
Useful6/10
Difficulty6/10
Novelty6/10
Unverified
2026
Replace an unconstrained recurrent transition with block-diagonal planar rotations whose angles are learned or conditioned on a slowly varying context variable. The resulting hidden-state norm and each two-dimensional block energy are exactly invariant in the ideal recurrence, preventing exploding or vanishing recurrent dynamics while retaining phase information over long horizons.
Useful6/10
Difficulty4/10
Novelty4/10
Unverified
2026
Replace repeated substrings in long sequences with nonterminal symbols from an acyclic straight-line grammar, then run the transformer on the compressed sequence. Unlike ordinary fixed tokenization, the compression objective explicitly minimizes the number of reusable binary productions, allowing repeated document-specific or corpus-level motifs to become single units. An expansion map lets the model recover token-level outputs for selected positions.
Useful6/10
Difficulty6/10
Novelty6/10
Unverified
2026
Construct a recurrent or state-space layer whose equilibrium Jacobian is placed near a nondegenerate Bogdanov–Takens point, then use a small unfolding parameter to move between damped, oscillatory, and slowly relaxing regimes. Unlike eigenvalue-only initialization near one, this controls both the double-zero center structure and the quadratic nonlinear coefficients that determine the local phase portrait.
Useful6/10
Difficulty6/10
Novelty8/10
Unverified
2026
Represent hierarchical routing decisions by one weighted bipartite graph between parent experts or regions and fine cells or token groups. Select a regularized edge set once, then derive both coarse parent activation and fine-grained routing from it, preventing later refinement from invalidating earlier load balancing.
Useful6/10
Difficulty6/10
Novelty6/10
Unverified
2026
Replace ordinary stride-2 pooling by a stochastic block-to-center map that is equivariant under global sign reversal and monotone in every input spin. For a binary feature channel, the layer computes the probability of a positive coarse feature from the number of positive fine features, samples or relaxes the resulting Bernoulli variable, and learns only a constrained scalar rather than an unconstrained pooling kernel. The same construction can be applied independently to channels or to graph…
Useful6/10
Difficulty4/10
Novelty6/10
Unverified
2026
Construct a sparse fixed orthogonal mixer by repeatedly applying pi/4 rotations to randomly matched pairs of feature coordinates. Place this mixer before top-k feature pruning, sparse projection, or activation quantization so that information is spread across coordinates without using a dense random matrix.
Useful6/10
Difficulty4/10
Novelty6/10
Unverified
2026
Replace a trainable shallow hidden layer by a deterministic feature dictionary generated from Chebyshev-spaced scalar parameters and quasi-uniform ridge directions. Train only the output linear map, or use the frozen layer as the first stage of a larger network, thereby eliminating hidden-layer backpropagation while retaining a constructive smooth-function approximation guarantee.
Useful6/10
Difficulty4/10
Novelty6/10