Architecture ideas

Research ideas extracted from mathematics papers, categorized as Architecture.

Unverified 2026

KS-Balanced Spectral Residual Block

Replace or augment a residual neural layer with a Fourier-domain scale-selective flow containing a learned second-order term and a fourth-order stabilizer. The block permits controlled low-frequency amplification, as required by the KS infrared mechanism, while damping high-frequency feature noise and preventing unbounded spectral growth.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Large scale behavior in the Kuramoto-Sivashinsky equation: The Schwinger-Dyson route arXiv:2607.17915
Unverified 2026

Anisotropic Anharmonic Diffusion Layer

Insert a learnable semigroup layer that evolves features according to a positive operator combining frequency damping and spatially varying confinement. Unlike isotropic Gaussian smoothing, the layer can damp selected frequencies differently along different axes and can suppress activations in learned spatial regions, while the positive-semigroup construction prevents amplification.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Sharp Time-Decay Estimates for Fractional Heat Semigroups Associated with Polynomial Anharmonic Oscillators arXiv:2607.17580
Unverified 2026

Parity-Constrained Signed Propagation

Replace the unsigned adjacency used by a deep message-passing network with a signing selected from an affine family that makes designated short even cycles unbalanced. Search this family for a small even-power trace, which acts as a proxy for a smaller spectral radius and suppresses explosive long-range propagation. The signing can be fixed before training, so the method adds no per-example inference cost.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Parity families and a kernel-averaged L-function for near-Ramanujan signings arXiv:2607.17343
Unverified 2026

Bounded-Width Neighborhood Signature Compression

Replace a dense node-to-landmark graph-attention or message-passing relation by a dictionary of distinct landmark-neighborhood signatures. Nodes sharing the same signature reuse the same structural landmark aggregate, while their individual hidden states are still passed through the output MLP, preserving node-specific predictions. On bounded-treewidth graphs the number of distinct signatures is provably linear in the number k of landmarks, with an explicit dependence on treewidth.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Neighbourhood complexity and identification problems for graphs of bounded treewidth and pathwidth arXiv:2607.16889
Unverified 2026

Regular Linear Hypergraph Attention

Construct attention groups as hyperedges of a linear r-uniform hypergraph: every pair of tokens is allowed to share at most one group, while each token participates in approximately the same number of groups. Apply local attention inside each group and aggregate the outputs across groups. The construction inherits the paper's sharp capacity bound and prevents both redundant pair interactions and high-degree token hubs.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Linear Turán Numbers of Uniform Hypertrees arXiv:2607.16854
Unverified 2026

Cycle-certified nonnegative rank-one attention

Represent a nonnegative attention or routing score matrix by two nonnegative vectors, X = uv^T, and learn only entries on a sparse bipartite graph of important query-key or token-expert interactions. Complete the remaining entries multiplicatively and monitor cycle residuals as a certificate of whether the sparse representation is compatible with rank one. Use local ratio violations to trigger additional edges or relax the rank-one approximation only where needed.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Tight Conic Relaxations for Rank-one Doubly Nonnegative Matrix Completion arXiv:2607.16796
Unverified 2026

Concave Heterogeneous Prototype Layer

Replace ordinary k-means-style prototype assignment with a distance-decay capture layer whose scale varies across samples, tokens, or classes. Train with a cooperative concave surrogate over prototype centers and anneal toward hard nearest-prototype assignment; this explicitly preserves useful gradients for multiple nearby prototypes while retaining sparse facility-like behavior at inference.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: When Is Heterogeneous Distance-Decay Facility Location Tractable? A Structural Classification, Exact Methods, and a Real-World Study arXiv:2607.16764
Unverified 2026

Dyson Diagonal Scaling for Directed Message Passing

Replace ordinary row-degree or symmetric normalization in a directed graph neural network with a nonlinear Dyson scaling. For a nonnegative directed adjacency matrix A, solve a positive vector equation and propagate with B = D A D, where D is the diagonal matrix of the solution. The resulting operator has row sums strictly below one, giving an explicit bound against exploding directed message propagation while retaining asymmetric edge information.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Non-symmetric vector dyson equations arXiv:2607.16333
Unverified 2026

Gray-coded numeric embeddings

Replace a learned embedding for each integer with a compositional embedding of its variable-length Gray codeword. Add an auxiliary constraint that numerically adjacent values have nearby representations, while preserving the ordinary task loss so that the model can learn when numerical adjacency matters.

Useful6/10
Difficulty4/10
Novelty8/10
Paper: Variable-length Gray codes for the Natural Numbers arXiv:2607.16088
Unverified 2026

Sparse Root-of-Unity Isometric Mixer

Replace a dense channel-mixing matrix by a sparse complex generalised weighing matrix W with exactly w nonzero entries in every row and column, then use U=W divided by square root of w as a norm-preserving mixer. Restricting to k=2 gives a real matrix with entries in {+1,-1}; k=4 supports signed phase rotations. The exact isometry should preserve signal and gradient norms while reducing channel-mixing cost.

Useful6/10
Difficulty6/10
Novelty7/10
Paper: Complex generalised weighing matrices in centraliser algebras of monomial representations arXiv:2607.16069
Unverified 2026

Rayleigh-Jeans Condensing Router

Replace a standard softmax MoE router with a thermodynamic router whose expert occupations maximize entropy subject to a prescribed total routing mass and mean routing energy. At high temperature, traffic is distributed across many experts; as temperature decreases or the energy budget tightens, traffic undergoes a predictable condensation transition in which excess load moves to the lowest-energy expert or expert group. This supplies an explicit control knob for adaptive specialization instead…

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Thermodynamic theory of voting and EU elections arXiv:2607.15119
Unverified 2026

RLCT-Certified Singular Attention

Use an anisotropic singular relative-position kernel in attention or graph message passing, with its exponent constrained by the paper's local integrability threshold. The module can represent sharper directional interactions than an RBF while providing an explicit certificate that its spatial gradient belongs to a chosen L^p space.

Useful6/10
Difficulty5/10
Novelty9/10
Paper: Geometric Criteria for Morrey Admissibility via the Real Log-Canonical Threshold arXiv:2607.14991
Unverified 2026

DynaBase Retrieval Forecast Head

Replace a parameter-heavy recurrent transition, or use this as a fallback, with a two-parameter nearest-neighbor successor blend in latent space. Given a query latent state, retrieve the closest state from an in-context trajectory and combine the query, the retrieved state, and its observed successor; this gives a zero-shot dynamical forecast with almost no trainable transition parameters.

Useful6/10
Difficulty4/10
Novelty5/10
Paper: A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems arXiv:2607.14937
Unverified 2026

Line-Graph Spectral Edge Parameterization

Replace one independently learned vector per graph edge with a truncated spectral expansion on the line graph. The model learns coefficients for low-frequency edge modes and reconstructs edge features before message passing, reducing parameters while imposing an inductive bias that incident edges should have correlated behavior.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Lossy compression of weighted graph adjacency matrices by transform coding arXiv:2607.14834
Unverified 2026

Moment-Resolved Stochastic Reservoir Readout

Replace mean-only readout from a noisy recurrent or Langevin reservoir by concatenating empirical first, second, and fourth raw moments of each hidden coordinate. The second and fourth moments retain input-dependent width and tail information generated by nonlinear confinement, while multiple independently initialized reservoirs can be concatenated before the final linear classifier to preserve complementary features.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Moment-Resolved Readout and Reservoir Diversity in Nonequilibrium Langevin Computing arXiv:2607.14520
Unverified 2026

Second-Order Sinh-Gordon Implicit Layer

Insert a differentiable implicit layer that maps boundary features to an interior latent field by solving a discrete sinh-Gordon equation. The paper's second-order convergence result motivates using a symmetric five-point discretization and a damped Newton solve rather than asking a neural network to learn the entire interior field directly.

Useful6/10
Difficulty6/10
Novelty7/10
Paper: Approximation of solutions of the sinh-Gordon equation $Δu -\sinh(2u)=0$ by hyperbolic orthogonal ring patterns arXiv:2607.14348
Unverified 2026

Threshold-Projection Recurrent Memory

Replace or augment a recurrent cell with multiple hysteresis memory branches whose states remain unchanged while the input stays within a branch-specific radius, then move toward the current input only when that radius is exceeded. The resulting cell has explicit persistence and bounded state changes, giving it an inductive bias for temporal hysteresis and reducing the need for the network to learn long-term memory behavior from scratch.

Useful6/10
Difficulty4/10
Novelty7/10
Paper: Accounting for Hysteresis and Eddy Currents in Finite Element Simulations of Ferromagnetic Laminated Cores using a Recurrent Neural Network arXiv:2607.14321
Unverified 2026

Spectral latent geometry for sparse attention

Build a sparse graph by thresholding normalized token or item inner products, then use the leading eigenvectors of its centered adjacency matrix as geometric features or a low-rank attention-logit bias. The graph avoids storing all pairwise similarities, while the paper's spectral bound supplies a concrete signal-to-noise test for deciding whether the resulting embedding is trustworthy.

Useful6/10
Difficulty5/10
Novelty6/10
Paper: Spectral Concentration and Recovery in Sparse High-Dimensional Random Geometric Graphs arXiv:2607.14304
Unverified 2026

Capacity-aware compressive-plus-indexed memory

Replace a purely recurrent or state-space history summary with two explicitly separated paths: a fixed-size state channel for compressed sequence mixing and a query-dependent indexed channel for exact or near-exact retrieval. Train a lightweight gate to invoke top-k retrieval only when the recurrent state has insufficient evidence for the current query, preserving near-constant cost on ordinary tokens while preventing catastrophic failures on long-range exact-recall tasks.

Useful6/10
Difficulty5/10
Novelty4/10
Paper: The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale arXiv:2607.14144
Unverified 2026

Gaussian Simplex Classification Head

Replace the unconstrained final classifier with equal-norm regular-simplex class directions and train it under explicit isotropic Gaussian feature noise. At fixed signal energy and equal class priors, the paper's Gaussian-max theorem predicts that this geometry maximizes finite-noise maximum-likelihood decoding probability, making it a concrete candidate for robust classification heads.

Useful6/10
Difficulty4/10
Novelty4/10
Paper: Stochastic Domination of Gaussian Maxima: A Resolution of the Weak Simplex Conjecture arXiv:2607.14087
Unverified 2026

Signature-memory neural CDE

Replace an unconstrained recurrent memory with a truncated path-signature state that is updated continuously from the input control path. Feed this structured state to a learned vector field, allowing the model to represent path-dependent dynamics through iterated integrals of the entire history rather than only the latest hidden state.

Useful6/10
Difficulty5/10
Novelty5/10
Paper: Dynamic Universal Approximation via Signature Controlled Differential Equations arXiv:2607.13886
Unverified 2026

Hermite-Schatten spectral layer

Replace a dense learned linear operator on continuous or image features by a truncated Hermite projection expansion whose coefficients are directly regularized in a Schatten-p norm. The layer becomes a structured low-rank operator, while the radial Hermite-Laguerre correspondence provides an analytically tractable parameterization and an exact spectral penalty.

Useful6/10
Difficulty6/10
Novelty7/10
Paper: Quantitative Fourier Restriction Estimates for Weyl Operators: Fourier-Support Dependence and Lower Bounds arXiv:2607.13697
Unverified 2026

Deadline-Aware Fair-to-Greedy Router

Use deadline objectives to train or control a router that explicitly trades off completion probability against completed work by a fixed horizon. Begin with fair allocation for robust exploration, then anneal toward a feedback-greedy rule once per-item difficulty estimates have sufficient evidence.

Useful6/10
Difficulty6/10
Novelty7/10
Paper: Meeting Uncertain Threats with Feedback arXiv:2607.13648
Unverified 2026

Finite-Orbit Circulant State Core

Replace the linear state transition in a small recurrent or state-space module by a circulant matrix acting on a vector over a finite field. The hidden state then has only finitely many possible values and follows an exactly periodic orbit after at most \(q^n\) states, eliminating numerical drift on modular-counting and symbolic-memory tasks. A learned real-valued encoder and decoder can surround the discrete core, while the transition itself is fixed, searched, or trained with a…

Useful6/10
Difficulty6/10
Novelty8/10
Paper: Periodicities in the Riordan arrays of polynomials over finite fields arXiv:2607.13442