The sharp reverse Hardy inequality in BMO for nonincreasing functions
arXiv:2608.08093
2026
Architecture
1 ideas extracted · analyzed Sep 1, 2026
What the math gives to ML
The paper proves a sharp lower bound for the Hardy prefix-averaging operator on the cone of nonincreasing functions: averaging cannot reduce the BMO seminorm by more than the optimal factor alpha_0, and one-jump functions are extremizers. The transferable asset is not generic averaging, but a certified anti-collapse property obtained by combining prefix integration with a monotonicity constraint. This suggests a causal neural module that applies cumulative averaging to monotone gates or feature profiles while monitoring a discrete BMO ratio. The construction could provide a cheap alternative to dense causal attention when importance profiles are expected to decay with sequence position.
Ideas from this paper
Unverified
2026
Replace a learned causal mixing profile by a monotone profile followed by a prefix-average Hardy mixer. The monotonicity constraint makes the mixer provably non-degenerate in the BMO sense: localized variation in the profile cannot be reduced below a calibrated factor by prefix averaging. This is a cheap alternative to dense causal attention for tasks where importance or state profiles are expected to decay along sequence position.
Useful5/10
Difficulty4/10
Novelty8/10