Architecture ideas

Attention variants, state-space and recurrent cells, normalization and token-mixing schemes — each tested against the standard block it replaces.

Unverified 2026

Imaginary-Axis Gramian Compression for Neural SSMs

Replace a large stable linear state-space or recurrent layer by a lower-order balanced realization computed from frequency-targeted controllability and observability Gramians. Use generalized low-rank ADI with imaginary-axis shifts concentrated at frequencies that dominate the training data, then retain states associated with the largest approximate Hankel singular values. This should reduce recurrent inference cost while preserving the layer's input-output response in the selected frequency…

Useful6/10
Difficulty6/10
Novelty7/10
Paper: A New Low-Rank Cholesky-Factor ADI Algorithm Allowing Shifts Anywhere in the Complex Plane with Applications to Data-Driven Model Reduction arXiv:2607.21969
Unverified 2026

Noncommutative adjacency-degree moments

Augment a graph neural network with features generated by noncommutative words in the adjacency matrix and diagonal degree matrix. Ordered patterns such as AD^2A and DADA distinguish where degree information occurs along a walk; the paper proves that the full scalar moment family determines every tree.

Useful6/10
Difficulty4/10
Novelty6/10
Paper: Adjacency-degree algebras and spectral determination of graphs arXiv:2607.21494
Unverified 2026

Approximation-Aware Hard-Core Routing

Construct a sparse routing or graph-neural architecture whose activation gates satisfy a hard-core constraint: neighboring sites, experts, or token groups cannot be active simultaneously. Compare the same local routing rule on bipartite and random regular interaction graphs; the graph structure should change the maximum usable activation dimension and may also change optimization stability.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: A Hard-Core Subshift Whose Sofic Mean Dimension Depends on the Sofic Approximation arXiv:2607.21398
Unverified 2026

Entropy-Calibrated Non-Backtracking Message Passing

Replace ordinary graph propagation, which repeatedly revisits the edge it just traversed, with a directed-edge non-backtracking operator. Normalize its learned gain using an estimate of the Hashimoto spectral radius so that feature magnitudes neither explode on high-growth graphs nor vanish on sparse graphs.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Critical-exponent spectra and rank two inverse realization on biregular trees arXiv:2607.21294
Unverified 2026

Linear-solve ensemble controller

Add a shallow neural interpolation controller to a neural ODE or state-space model so one shared vector field matches prescribed derivatives at several anchor trajectories. At every control time, compute controller weights from a small linear system instead of learning all task-specific parameters by backpropagation.

Useful6/10
Difficulty5/10
Novelty5/10
Paper: Exact ensemble controllability for neural differential equations via neural interpolation arXiv:2607.21112
Unverified 2026

Multiplicative Adaptive Attention Graph

Give each query-token pair a positive adaptive edge weight that evolves by a multiplicative rule instead of relying only on instantaneous dot-product attention logits. Edges whose aggregate interaction is useful can grow, while overloaded or incompatible neighborhoods can shrink. Sparse initialization is preserved because an edge initialized at zero remains zero under the multiplicative dynamics.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: The mean-field limit of non-exchangeable particle systems with non-conservative dynamics and adaptive weights arXiv:2607.21110
Unverified 2026

Reflected Survival Routing

Replace independent binary early-exit or token-pruning decisions with a monotone randomized survival process for each token or expert route. A token can lose survival mass at each layer but cannot become active again; the model is trained with a reflected obstacle-style penalty that activates when the predicted value of continuing computation is below the value of stopping plus the compute cost. Mean-field statistics are computed over currently surviving tokens, making routing less sensitive to…

Useful6/10
Difficulty5/10
Novelty5/10
Paper: A new probabilistic approach for mean field games of optimal stopping arXiv:2607.21062
Unverified 2026

L2-Certified DAG Attention Ordering

Add a learned scalar ordering to a directed graph attention layer and retain only forward edges, producing a DAG attention mask without requiring a supplied topological order. Train the ordering with a differentiable surrogate for weighted surplus, and regularize it toward the paper's explicit half-weight-minus-l2 certificate. This supplies a principled alternative to random masking or unconstrained bidirectional graph attention when causal or hierarchical information flow is desirable.

Useful6/10
Difficulty5/10
Novelty8/10
Paper: The optimal constant for minimum weight feedback arc sets in oriented graphs arXiv:2607.20996
Unverified 2026

Degree-Capped Simplicial Residual Step

Set the residual propagation coefficient of a simplicial neural layer from a cheap upper bound on the operator spectrum instead of tuning it blindly. The degree-majorization theorem supplies a bound on the largest eigenvalue, while the Brouwer-type inequality supplies a topology-count-based bound on sums of the top eigenvalues.

Useful6/10
Difficulty3/10
Novelty6/10
Paper: Degree Majorization and Laplacian Eigenvalue Sums for Simplicial Complexes arXiv:2607.20910
Unverified 2026

Separable Tensor-Product Spline Trial Layer

Use the paper's correspondence between KAN splines and finite-element or isogeometric shape functions to build coordinate-separable tensor-product trial layers. Replace additive coordinate aggregation with a multiplicative contraction of one-dimensional spline expansions, yielding an explicit tensor-product basis without storing a dense multidimensional grid.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: PG-KINN: A Physics-Informed Petrov-Galerkin Kolmogorov-Arnold Network for Solving Forward and Inverse PDEs arXiv:2607.20378
Unverified 2026

Win-Martingale Adaptive Router

Replace a conventional softmax router or fixed halting score with a scalar confidence state that evolves as a bounded martingale diffusion. The state starts at the network's prior confidence, receives evidence-dependent stochastic increments, and is absorbed at 0 or 1; absorption selects an MoE expert or halts additional transformer blocks. State-dependent volatility lets the model explore aggressively when uncertain and commit rapidly when confident, while the martingale constraint prevents…

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Embedding martingale diffusions as binary posteriors in sequential inference arXiv:2607.20373
Unverified 2026

Constraint-Free Skew Coupling

Compose independently parameterized neural dynamical modules through power-preserving skew coupling instead of equality penalties or projected constraints. This creates a modular graph or world model in which information exchanged between modules is antisymmetric, so internal coupling cannot create or destroy total latent energy.

Useful6/10
Difficulty6/10
Novelty6/10
Paper: Mixed finite element discretization of intrinsic geometrically exact beams for explicit multibody dynamics arXiv:2607.20245
Unverified 2026

Strongly Pseudomonotone Implicit Router

Replace an explicit MoE router or constrained output head with the solution of a variational inequality over a convex feasible set. The neural operator can be nonmonotone, but training should enforce a measurable strong-pseudomonotonicity margin so the selected route or control is unique and has bounded sensitivity to changes in the token representation. Use an explicit projection residual for approximate solving and for monitoring whether the implicit layer has actually converged.

Useful6/10
Difficulty6/10
Novelty5/10
Paper: A Coupled Nonsmooth Dynamical System: Global Well-Posedness, Stability and Sensitivity Analysis arXiv:2607.20133
Unverified 2026

Nestohedral Adaptive Token Tree

Replace fixed sequence-to-sequence attention with a dynamically maintained tree of connected token groups. Groups can be fused to reduce the number of attention units or split when their representation is heterogeneous, while hypergraph connectivity and nestedness ensure that every intermediate hierarchy remains valid.

Useful6/10
Difficulty6/10
Novelty5/10
Paper: Generalised flip order on the faces of nestohedra arXiv:2607.20132
Unverified 2026

Grassmannian Tropical Router

Replace unconstrained MoE router logits with structured phase scores indexed by N-subsets of M ordered parameters. Each token is assigned to the dominant phase, while neighboring routing regions obey the Grassmannian rule that adjacent labels share N-1 indices, reducing arbitrary fragmented decision boundaries and encouraging smooth expert transitions.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Combinatorial geometry of the 2D Toda lattice and Davey Stewartson equation arXiv:2607.20109
Unverified 2026

Reactive Mass-Weighted Message Passing

Augment every graph or set token with a positive learned mass M_i that controls how strongly it contributes to other nodes and evolves through a growth-minus-inhibition equation. Use separate learned interaction kernels for state transport and mass inhibition, while retaining a directed interaction matrix so the layer is not forced to be permutation-symmetric or conservative.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: A note on application of mean-field limit to non-exchangeable non-conservative systems arXiv:2607.20014
Unverified 2026

Boolean Tree Token Routing

Replace unconstrained pairwise token grouping with a tree whose edges carry independent merge or cut variables. The connected components of the retained edges define a valid partition at every forward pass, while learned edge gates control the amount of token aggregation. A coarse component-level computation can then replace part of dense attention.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: On Boolean sublattices of finite partition lattices arXiv:2607.19940
Unverified 2026

Hierarchical B-spline sparse-grid front end

Replace a dense tensor-product positional encoding or first MLP layer with a hierarchical sparse-grid B-spline feature map. The network evaluates only localized basis functions indexed by multi-levels with bounded total level, reducing feature count while retaining high-order approximation for functions with mixed derivative regularity. The basis can initially be fixed and later fine-tuned jointly with the downstream network.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: A hierarchical sparse-grid particle method for the Vlasov--Poisson system arXiv:2607.19898
Unverified 2026

Hemifield sum-difference orientation channels

Compute separate doubled-angle orientation order parameters for left and right image regions, then expose their sum and difference as symmetric and antisymmetric global features. This gives a network a low-dimensional inductive bias for global vertical structure versus left-right imbalance, while retaining magnitude channels that indicate when either readout is undefined because orientations cancel.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Perceived vertical and eye level as one orientation order parameter: a closed-form account of the Li-Matin rules for egocentric space arXiv:2607.19681
Unverified 2026

Hashed Local-Density Particle Layer

Replace an O(N^2) kernel-density interaction in a particle neural SDE or diffusion sampler with a clipped, randomly shifted histogram density estimate. Feed the local estimated density into the particle drift as a multiplicative gain, preserving density-dependent dynamics while evaluating all particles through occupied-cell hashing in expected O(N) time for fixed dimension and number of shifts.

Useful6/10
Difficulty5/10
Novelty8/10
Paper: Density-Dependent McKean--Vlasov Diffusions: Subgaussian Occupancy Bounds and Polynomial Propagation of Chaos arXiv:2607.19583
Unverified 2026

Noise-Aware Soft ECOC Decoding

Treat the binary outputs of the hyperplane head as a noisy channel and decode with reliability-weighted likelihood rather than unweighted Hamming distance. Estimate each bit's flip probability on validation data and give unreliable hyperplanes less influence, while retaining the logarithmic code-length scaling.

Useful6/10
Difficulty3/10
Novelty6/10
Paper: Fundamental limits of distributed multiclass classification from simple binary decisions arXiv:2607.19334
Unverified 2026

Logarithmic Random-Hyperplane Classifier

Replace a $K$-class softmax with $N$ binary hyperplane heads, where each class is represented by the signs of its projections onto fixed random directions. Train the embedding to reproduce these codewords and decode by nearest Hamming codeword. The paper's guarantee suggests that $N\approx 2\log_2 K+\log_2(1/\delta)$ can separate all class centers with high probability in sufficiently high dimension, giving a concrete width rule rather than choosing the number of binary heads heuristically.

Useful6/10
Difficulty4/10
Novelty5/10
Paper: Fundamental limits of distributed multiclass classification from simple binary decisions arXiv:2607.19334
Unverified 2026

Slepian-Concentration Feature Initialization

Initialize a coordinate-network feature bank with the leading eigenfunctions of a bandlimited concentration operator instead of random Fourier features. For a desired spatial region E, these features maximize the fraction of their L2 energy inside E among all functions with frequency support in Omega, giving a principled basis for localized signals.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Optimal concentration in the Paley-Wiener space arXiv:2607.19192
Unverified 2026

Mollified Transport-Quantile Layer

Use a transport map \(Q_\theta\) from a fixed latent reference distribution to a data distribution, but expose only its locally averaged version \(\bar Q_{\theta,\sigma}(z)=\mathbb E_{u\sim K_\sigma(\cdot-z)}Q_\theta(u)\). Latent-space mollification integrates the pole-type influence singularity instead of allowing one training sample near \(Q_\theta(z)\) to dominate the quantile feature or its gradient.

Useful6/10
Difficulty4/10
Novelty7/10
Paper: The Influence Function of Transport-based Quantiles arXiv:2607.19080