The geometry of AI validation: Exact certification limits for iid best-of-N search
arXiv:2608.21496
2026
Theory
1 ideas extracted · analyzed Sep 1, 2026
What the math gives to ML
This paper gives an exact, constructive limit on what reliability audits can certify for iid best-of-N selection when audits only measure smaller search budgets. The key transferable asset is an explicit ambiguity width, B_{m,N}, showing that repeated evaluation at n <= m cannot distinguish reliability worlds that differ substantially at deployment budget N; the relevant scale is m^2/N. This can become an audit-aware validation layer for best-of-N decoding, reward-model selection, or self-consistency, preventing false claims of reliability when the validation budget is structurally too small. It also provides a principled way to choose audit budgets rather than relying on ad hoc numbers of sampled candidates.
Ideas from this paper
Unverified
2026
Attach an exact structural uncertainty report to any best-of-N evaluation: after measuring reliability only for budgets n = 1,...,m, report that deployment reliability at budget N is unresolved by at least B_{m,N}. Use this width to select the smallest audit budget that makes a claimed reliability gap meaningful, or reject model comparisons whose validation budget lies below the square-root-of-N threshold.
Useful6/10
Difficulty3/10
Novelty8/10