Asymptotic bias of the plug-in Shannon entropy estimator under a regularly varying occupancy model
arXiv:2607.27721
2026
Regularization
1 ideas extracted · analyzed Aug 31, 2026
What the math gives to ML
The paper gives a sharp finite-sample bias law for the plug-in Shannon entropy estimator when the underlying categorical probabilities have a regularly varying heavy tail. This is transferable to sampled-policy entropy bonuses, token-distribution diagnostics, and entropy-based regularizers, where minibatch counts systematically underestimate entropy because rare categories are unseen. The practical adaptation is an occupancy-based correction: estimate the tail index from distinct-category growth and replace the unknown asymptotic factor using the observed number of distinct categories. This should be tested against uncorrected plug-in entropy and standard small-sample corrections on heavy-tailed categorical distributions.
Ideas from this paper
Unverified
2026
Correct minibatch or trajectory-based categorical entropy estimates using the paper's power-law occupancy asymptotic. The corrected estimate adds back entropy lost through unseen rare categories, with the correction magnitude inferred from the number of distinct observed categories and an estimated tail index.
Useful6/10
Difficulty5/10
Novelty7/10