Asymptotic bias of the plug-in Shannon entropy estimator under a regularly varying occupancy model

arXiv:2607.27721 2026 Regularization 1 ideas extracted · analyzed Aug 31, 2026

What the math gives to ML

The paper gives a sharp finite-sample bias law for the plug-in Shannon entropy estimator when the underlying categorical probabilities have a regularly varying heavy tail. This is transferable to sampled-policy entropy bonuses, token-distribution diagnostics, and entropy-based regularizers, where minibatch counts systematically underestimate entropy because rare categories are unseen. The practical adaptation is an occupancy-based correction: estimate the tail index from distinct-category growth and replace the unknown asymptotic factor using the observed number of distinct categories. This should be tested against uncorrected plug-in entropy and standard small-sample corrections on heavy-tailed categorical distributions.

Ideas from this paper

Unverified 2026

Regular-Variation Entropy Debiasing

Correct minibatch or trajectory-based categorical entropy estimates using the paper's power-law occupancy asymptotic. The corrected estimate adds back entropy lost through unseen rare categories, with the correction magnitude inferred from the number of distinct observed categories and an estimated tail index.

Useful6/10
Difficulty5/10
Novelty7/10
Paper: Asymptotic bias of the plug-in Shannon entropy estimator under a regularly varying occupancy model arXiv:2607.27721