Unverified
2026
Attach two oscillator channels to each recurrent, state-space, or graph hidden unit and convert them into a phase field over nodes or spatial positions. Encode every overlapping triple of neighboring phases as one of the 13 weak ordinal patterns, including seven near-tie patterns, then use the resulting normalized entropy and pattern frequencies to detect hidden-state collapse, coherent clustering, or transient regime changes. During training, either use the entropy only as a controller for…
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Construct a contractive multi-branch recurrent or generative network whose branches define an iterated-function system, and regularize it so that branch entropy is high relative to average contraction while compositions remain exponentially separated. The target is a measurable attractor-dimension law rather than only a benchmark improvement: the invariant measure dimension should approach min(d, H divided by chi), where d is state dimension.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace overflow dropping in a mixture-of-experts layer with a discrete-convex load repair procedure. The router first chooses experts from neural logits, then applies capacity-aware exchange moves that preserve the total number of dispatched tokens and monotonically improve the routing objective whenever a feasible swap exists.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace an unconstrained message-passing or recurrent propagation matrix by a directed-edge operator with non-backtracking connectivity and orientation-dependent turning phases, inspired by the Kac–Ward construction. During training, monitor and control the zero-momentum spectral gap of \(\mathcal A(0)=I-K(0)\), keeping the model near but on the stable side of the critical surface to obtain long memory without uncontrolled amplification.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Augment a neural-network update with an auxiliary, damped stochastic branch that acts like the paper's floating dissipative reservoir. A trainable mixing phase \(\phi\) combines the task-gradient branch and auxiliary branch; \(\phi\) is adapted to make the auxiliary response to a chosen control perturbation nearly zero while retaining a finite task-gradient response. The intended benefit is selective insensitivity to nuisance hyperparameters or perturbations, with a measurable response peak…
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Add valid-inequality penalties to a segmentation model that predicts both node cut probabilities and pairwise separation probabilities. The penalties enforce that a predicted pair cannot be separated without an appropriate vertex separator, and that local path and intersection relations among pair predictions remain feasible. This supplies structural supervision even when only sparse or noisy pair labels are available.
Useful6/10
Difficulty4/10
Novelty6/10
Unverified
2026
Represent hidden features using a tensor-product polynomial-evaluation code instead of storing one value per feature. Corrupted coordinates can then be identified through violations of low-degree consistency and repaired before the next neural layer, targeting robustness to hardware faults, unreliable memory, malicious distributed workers, and adversarial activation corruption.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace an unconstrained multiplicative interaction between two nonnegative neural features by a lifted gate whose first and second moments satisfy the paper's semidefinite relaxation for the set F = {(x1,x2): x1,x2 >= 0, x1 x2 <= 1}. Insert the gate into an MLP, attention score, or MoE router to prevent explosive feature products while retaining a tractable convex feasible set.
Useful6/10
Difficulty7/10
Novelty8/10
Unverified
2026
Insert a few implicit DLSS diffusion steps after a network produces a nonnegative spatial probability field, such as a segmentation map, density estimate, or normalized image likelihood. The layer is a nonlinear fourth-order smoother that preserves positivity and is contractive in square-root/Hellinger distance, potentially reducing prediction noise without ordinary Euclidean blurring.
Useful6/10
Difficulty7/10
Novelty7/10
Unverified
2026
Partition a neural state or feature vector into blocks and identify directed dependencies between blocks from one-step transition data. Use the inferred design structure matrix as a hard mask or soft gate on recurrent, state-space, graph, or mixture-of-experts couplings, replacing a dense unconstrained interaction matrix with a data-supported sparse graph.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Augment node features with eigenvectors corresponding to negative eigenvalues of the Bethe-Hessian H(t,G), rather than using only Laplacian or adjacency positional encodings. The diagonal D-I correction is designed for sparse, locally tree-like graphs and should suppress degree-fluctuation artifacts near the connectivity threshold. Feed the resulting coordinates to a GNN through a learned gate so the model can ignore them when they are uninformative.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Treat trainable prototypes, class centers, codebook entries, or router expert embeddings as interacting particles and add a mollified repulsive Coulomb force to their task-gradient update. Unlike a fixed repulsion coefficient, use the paper's explicit density envelope to reduce repulsion over training and use the associated density-dependent mollification radius, so early training prevents collapse while late training permits precise cluster formation.
Useful6/10
Difficulty4/10
Novelty5/10
Unverified
2026
Initialize a stable diagonal state-space layer with decay rates \(\omega_i=|\xi_i|\), where \(\xi_i\sim\mathcal N(\mu,\sigma^2)\), instead of using a narrowly clustered rate distribution. The nonzero density of rates near zero creates a population of slow modes whose aggregate impulse response has an algebraic tail, enabling long-horizon memory while every finite-dimensional mode remains exponentially stable.
Useful6/10
Difficulty4/10
Novelty6/10
Unverified
2026
Use two nested Riesz reconstruction spaces to estimate unresolved residual energy for every checkpoint. Under a measurable saturation assumption, convert the coarse and enriched monitors into lower and upper error bounds, and certify a unique checkpoint whenever its upper bound lies below every competitor's lower bound.
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Rank candidate prompts, demonstrations, critiques, or system instructions by how many bits of reproduction cost they save for a specified artifact distribution. Replace raw prompt-token heuristics with a paired score that rewards both higher success probability and lower generation computation, then train or retrieve prompts maximizing this score under a token budget.
Useful6/10
Difficulty5/10
Novelty6/10
Unverified
2026
Construct a neural residual block as a composition of positive-time flows from two learned vector fields, rather than one unconstrained residual update. Add a learned Lie-bracket correction channel so that the block can cancel leading noncommutative splitting errors without using negative coefficients. The resulting block has a tunable effective integration order while preserving forward-time behavior for dissipative dynamics.
Useful6/10
Difficulty7/10
Novelty7/10
Unverified
2026
Replace unconstrained or entropy-regularized MoE routing with a minimally disruptive update that preserves a lower bound on the log-determinant of the experts' weighted output span. The router still tracks the desired mixture, but a projection prevents the active experts from becoming linearly redundant or collapsing onto a low-rank subset.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Replace a dense multiresolution grid encoding for a coordinate MLP with a hierarchical sparse tensor-product encoding whose active cells are selected by local hierarchical surpluses. Combine anisotropic component grids with alternating binomial weights, then refine only regions whose encoded or prediction residual is large. This should preserve fine detail around localized structures while avoiding the exponential parameter count of a full grid.
Useful6/10
Difficulty5/10
Novelty5/10
Unverified
2026
Train a neural PDE solver using collocation points sampled from a fixed reference diffusion and a time weight that compensates for the point-start singularity. Replace the Euclidean Hessian by the intrinsic tensor Gθ=σD²uθσ, and use source Picard updates so that nonlinear curvature coupling is iterated under an explicit contraction target.
Useful6/10
Difficulty5/10
Novelty8/10
Unverified
2026
Treat a selected neural submodule as an open dynamical system embedded in the rest of the network. Regularize it to contain internal modes that are simultaneously reachable from many external features and observable through many external outputs, rather than behaving as a one-sided receiver, broadcaster, or disconnected read/write split.
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace an unconstrained entrywise nonlinearity on a positive Gram or covariance matrix by a learned scalar function satisfying the paper's finite-order positivity-preserver conditions. The transformed matrix remains PSD for matrices of the target width n, allowing nonlinear Gram propagation without eigenvalue clipping or projection.
Useful6/10
Difficulty5/10
Novelty7/10
Unverified
2026
Replace a conventional covariance or density-matrix discrepancy with the geodesic quantum f-divergence between an example's predicted positive-definite matrix and its target matrix. Use t as a controllable interpolation between the standard Petz divergence at t=0 and the maximal divergence at t=1, with f(x)=x log x or another operator-convex power generator. The loss is suited to covariance-predicting networks, SPD-valued embeddings, and matrix-valued classifiers where eigenvector alignment…
Useful6/10
Difficulty6/10
Novelty7/10
Unverified
2026
Replace Euclidean updates and interpolation of probability vectors in a mixture-of-experts router or attention simplex with updates in square-root coordinates, where the Fisher–Rao geometry is spherical. If the task has a desired neutral or calibrated family of distributions, represent that family as a linear subsphere in square-root space and project router outputs onto it after every update.
Useful6/10
Difficulty4/10
Novelty4/10
Unverified
2026
Parameterize a periodic neural vector field as the sum of a harmonic global drift, an exact gradient field, and a co-exact divergence-free field. This gives separate control over conservative attraction/repulsion, rotational transport, and domain-wide drift, potentially preventing one unconstrained MLP from entangling incompatible dynamics.
Useful6/10
Difficulty5/10
Novelty6/10