Architecture ideas

Research ideas extracted from mathematics papers, categorized as Architecture.

Unverified 2026

Schatten-Stable Noncommutative Matrix Layer

Replace an ordinary elementwise interaction between two feature matrices by a noncommutative functional-calculus layer \(\varphi(A,B)\), where \(A\) and \(B\) are Hermitian channel operators that need not commute. Add a soft penalty on \([A,B]=AB-BA\), and use a Besov-smooth parameterization of \(\varphi\) so that perturbations are controlled in Schatten \(p\)-norm for \(p\leq2\). This creates a principled matrix interaction module that can remain stable when feature operators or graph…

Useful5/10
Difficulty6/10
Novelty8/10
Paper: Commutator estimates for functions of noncommuting self-adjoint operators arXiv:2608.16731
Unverified 2026

Hermitian ETF classifier head

Construct a classifier whose normalized class vectors form an explicit 2d-line equiangular tight frame instead of using independently initialized weights. The ETF gives every class the same norm, equal pairwise coherence, and an isotropic frame operator, which should make final-layer gradients better conditioned and reduce accidental class crowding. The classifier can be fixed, or restricted to a learned unitary rotation of the ETF so that its geometry is preserved during training.

Useful5/10
Difficulty4/10
Novelty4/10
Paper: New constructions of optimal arrangements of $2d$ lines in $\mathbb{C}^d$ arXiv:2608.16116
Unverified 2026

PSD-Safe Learnable Similarity Kernel

Use the finite-order characterization to learn a nonlinear similarity function for token, patch, or graph-node Gram matrices while preserving PSD by construction or by a differentiable certificate loss. This creates a kernelized attention or graph-readout mechanism in which nonlinear affinity transformations cannot introduce indefinite similarity geometry.

Useful5/10
Difficulty6/10
Novelty6/10
Paper: A finite-order characterization of entrywise positivity preservers arXiv:2608.15904
Unverified 2026

Resolvent response basis for conditioned modules

Compress a module whose output changes with a scalar condition such as diffusion time, temperature, or compute budget by representing its response in a low-rank basis generated by resolvent-like functions. Distinct spectral modes produce rational factors \((1-\tau\lambda_k)^{-1}\), allowing a small number of learned components to approximate a large hypernetwork or condition-dependent parameter table.

Useful5/10
Difficulty6/10
Novelty7/10
Paper: Spectral duality structures and the Fisher--Rao geometry of reset distributions arXiv:2608.15805
Unverified 2026

Critical 2d Intensity Bottleneck

Replace a complex latent vector x in C^d by squared magnitudes of m learned complex linear projections. Set m equal to 2d: the paper proves that m less than or equal to 2d minus 1 cannot generically preserve the latent up to global phase, whereas m equal to 2d is generically sufficient, giving a principled minimal width for a phase-invariant neural bottleneck.

Useful5/10
Difficulty4/10
Novelty7/10
Paper: The Minimum Number of Measurements for Almost-Everywhere Complex Phase Retrieval arXiv:2608.15003
Unverified 2026

Correlated-Gaussian Orbit Fingerprint

Replace a polynomial layer's single-replica output statistics with a finite fingerprint computed from several correlated Gaussian replicas. Train the fingerprint to be invariant under orthogonal reparameterizations while remaining discriminative between genuinely different polynomial maps, preventing models from collapsing distinct tensor functions that have identical marginal output laws. This is a practical symmetry-aware regularizer or auxiliary embedding for tensorized MLPs and polynomial…

Useful5/10
Difficulty5/10
Novelty8/10
Paper: Finite Gaussian Reconstruction of Polynomial Orbits: From Correlated Moments to Oscillatory Periods arXiv:2608.14475
Unverified 2026

Analytic Gaussian Measurement Bottleneck

For a neural model whose outputs lie on a d-dimensional analytic family in a very high-dimensional space, replace the full output vector by 2d+1 or a modestly oversampled number of fixed Gaussian scalar measurements. The paper's theorem predicts almost-sure injectivity in the noiseless setting, so an inverse network or decoder can recover the same latent instance without processing the full observation. Because the theorem does not provide a noise-stability constant, use M=4d+8 or M=8d in the…

Useful5/10
Difficulty4/10
Novelty5/10
Paper: Analytic inverse problems with finitely many random measurements arXiv:2608.14324
Unverified 2026

Velocity-Scaled Symbolic Flow Model

Represent a continuous-time neural dynamical system as a symbolic Markov chain over regions together with a positive learned roof function giving the time spent in each region. Weight local reconstruction and prediction errors by the predicted vector-field speed, following the paper's scaled Hölder coding relation, so that the model does not over-penalize arbitrarily small coordinate errors near equilibria. This produces a hybrid latent model with discrete long-range structure and continuous…

Useful5/10
Difficulty6/10
Novelty7/10
Paper: Symbolic dynamics for non-uniformly hyperbolic flows arXiv:2608.14095
Unverified 2026

Jucys–Murphy spectral token mixer

Add a structured token-mixing layer based on commuting sums of swap operators rather than unconstrained pairwise attention. The layer learns a low-degree spectral filter in the Jucys–Murphy operators, allowing it to represent hierarchical interactions while retaining an explicit algebraic inductive bias.

Useful5/10
Difficulty6/10
Novelty7/10
Paper: Hitting-time mixing for the star transposition shuffle arXiv:2608.13727
Unverified 2026

Mutation-invariant interaction mixer

Represent a directed interaction graph by a Laurent-polynomial Euler-like matrix and use its evaluation as a signed message-passing or attention-mixing operator. During dynamic rewiring, require the new graph representation to preserve the associated bilinear form up to the congruence transformation induced by the change of basis, so equivalent routings produce equivalent hidden states.

Useful5/10
Difficulty6/10
Novelty8/10
Paper: Monodromy of plane curve singularities and quiver mutation arXiv:2608.12484
Unverified 2026

Pisot-Separated Multiscale Codes

Replace an unconstrained geometric multiscale codebook by features generated from a finite digit set and a Pisot scale factor. The contracting algebraic-conjugate directions should suppress near-collisions between representations at different scales, producing a discretely separated hierarchy that can be used for embeddings, recurrent memory, or quantized transformer states.

Useful5/10
Difficulty6/10
Novelty9/10
Paper: Self-similar Delone sets and Pisot numbers arXiv:2608.11867
Unverified 2026

Spectral Phase-Lifted Features

Augment a hidden representation with positively homogeneous interaction features built from approximate eigenmodes of a linear layer. Fractional products of mode magnitudes and phases provide nonlinear channels whose transformation laws are inherited from the spectrum of the underlying operator, potentially representing oscillatory or multiplicative dynamics more compactly than a generic MLP.

Useful5/10
Difficulty7/10
Novelty8/10
Paper: The spectrum of operator extensions to free Banach Lattices arXiv:2608.11437
Unverified 2026

Cross-Order Hodge Stability Scaling

Normalize every higher-order simplicial message-passing or diffusion block using the spectral radius of a lower-order up-Laplacian, rather than estimating a separate radius for each order. The paper's monotonicity theorem guarantees that this shared bound is conservative for all higher orders, enabling stable explicit updates with one spectral calibration.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Eigenvalue growth of the discrete Hodge Laplacian across dimensions arXiv:2608.11170
Unverified 2026

Birkhoff Modular Mask Regularizer

Represent structured neural masks or routing states as order ideals of a finite prerequisite poset, then use a modular score whose exact minimizers are a desired decomposition-closed family of valid configurations. This replaces many pairwise constraint penalties with one additive potential that gives zero cost to every intended valid state and positive cost to invalid intermediate states.

Useful5/10
Difficulty6/10
Novelty7/10
Paper: Decomposition-Closed Sublattices as Minimizer Sets of Modular Functions over Distributive Lattices arXiv:2608.10026
Unverified 2026

Lexicographic spectral activation

Replace an ordinary elementwise nonlinearity on a learned Hermitian matrix with a matrix function f(A), while supplying exact Jacobian-vector and Hessian-vector products through the lexicographic divided-difference formula. This gives a principled spectral layer for covariance features, graph operators, attention kernels, or matrix-valued embeddings, particularly when perturbation matrices do not commute and eigenvalues are repeated or nearly repeated.

Useful5/10
Difficulty5/10
Novelty5/10
Paper: Lexicographic functional calculus and its application to functional calculus calculus arXiv:2608.08404
Unverified 2026

Seven-Factor Stochastic Transition Layer

Parameterize a learned 3-state transition operator as a product of at most seven elementary row-stochastic matrices rather than learning its nine entries independently. Each factor performs one convex pull-in of row i toward row j, so every intermediate and final matrix remains row-stochastic and the layer has a sparse, bounded-depth interpretation.

Useful5/10
Difficulty4/10
Novelty7/10
Paper: Bang--bang representation of $3\times 3$ embeddable stochastic matrices arXiv:2608.08242
Unverified 2026

Rank-Collapse Quadratic Token Router

Replace independent token scores with a query-conditioned positive-semidefinite low-rank quadratic score over a fixed-size selected subset. Repeatedly convert the quadratic objective into a linear exposure vector and apply a cheap top-k oracle, allowing the selector to model joint token interactions without constructing an n-by-n attention matrix. The margin between the current low-dimensional shadow and alternatives provides a practical confidence or early-stopping signal.

Useful5/10
Difficulty5/10
Novelty6/10
Paper: The Rank-Collapse Principle for Quadratic Optimization arXiv:2608.07828
Unverified 2026

Sidon pair encoding

Assign each of K entity or token types an integer code from a B_{2,\Delta}-set A, so every unordered pair {i,j} produces a unique and margin-separated scalar code a_i+a_j. Use this code as a compact symmetric pair feature for graph edges, attention biases, or pairwise relation MLPs, avoiding collisions that occur when ordinary low-dimensional additive encodings are quantized or hashed.

Useful5/10
Difficulty4/10
Novelty7/10
Paper: Sidon sets with $Δ$-separated sumsets in additive number theory arXiv:2608.07416
Unverified 2026

Vacancy-preserving collision-free router

Build a differentiable assignment layer whose rows represent tokens and whose columns represent experts, memory slots, or attention slots. Each row has unit probability mass, but no column receives positive mass from two rows; maintaining at least one vacant column makes assignments continuously deformable through elementary vacancy moves instead of abrupt softmax switches.

Useful5/10
Difficulty6/10
Novelty5/10
Paper: Tilings, packings, and the existence of Schwartz-class Gabor windows arXiv:2608.06679
Unverified 2026

Squarefree Cycle Positional Encoding

Augment every graph node with weighted participation in simple cycles of lengths 3 through K, computed using the paper's squarefree trace construction. Feed these features into a graph transformer or message-passing network so nodes with identical local degrees and ordinary spectral statistics can still be distinguished by their exact loop environment.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Squarefree Matrix Formulas for the CWR Invariant of Alternating Knots and Links arXiv:2608.06372
Unverified 2026

Simultaneous-Label Sparse Attention

Replace a dense attention pattern by the exact intersection of a fixed or cheaply computed base graph H and a learned shared-label relation. Two tokens can exchange information only when they are adjacent in H and share at least one of d labels, producing a controllable structured sparsity pattern. The label count d becomes an explicit capacity and compute knob: increasing d enlarges the relation vocabulary without requiring a dense pairwise mask.

Useful5/10
Difficulty6/10
Novelty5/10
Paper: Simultaneous Graph Parameters and How to Bound Them arXiv:2608.06055
Unverified 2026

Fuzzy permutation attention

Replace part of an attention matrix with a mixture of fuzzy permutation matrices induced by short permutations. Each basis element represents an order-preserving k-token matching smeared over all embeddings into the sequence, while a balancing constraint makes the aggregate attention receive uniform global coverage. Retain a standard low-rank or local-attention residual so the structured branch does not prevent arbitrary content-dependent interactions.

Useful5/10
Difficulty5/10
Novelty7/10
Paper: Fuzzy latin squares and balanced permutation pattern statistics arXiv:2608.05335
Unverified 2026

Chain-Compatible Differential Pooling

Treat learned features on a mesh as differential forms and pool them against oriented chains using wedge or cap products instead of ordinary coordinate averaging. Couple forward and boundary features with the signed chain differential so that pooling commutes with differentiation, preserving local conservation and orientation information.

Useful5/10
Difficulty5/10
Novelty8/10
Paper: Differential Homology arXiv:2608.05048
Unverified 2026

Frame-safe totally-positive front-end

Replace the first learned one-dimensional convolution or STFT-like feature extractor with a differentiable bank of time-frequency shifts of a totally positive window. Parameterize the temporal spacing \(\alpha\) and frequency spacing \(\beta\) so that \(\alpha\beta<1\) is always satisfied, giving a mathematically certified oversampled representation instead of an arbitrarily subsampled filterbank.

Useful5/10
Difficulty5/10
Novelty5/10
Paper: Gabor Frames of Totally Positive Functions: A Complete Characterization arXiv:2608.04992