neural-collapse-is-forbidden-information-floors-in-language-models-833204ab·1 events·first seen Aliases: Neural Collapse Is Forbidden: Information Floors in Language Models
A new arXiv preprint argues that within-class variance in language model representations is not incomplete neural collapse but rather allocated information storage obeying a measurable law. Across 14 models spanning a 100x parameter range, macro-category structure accounts for only 4-12% of representational variance while within-token context carries 79-91%. The authors prove a theoretical floor on within-category dispersion proportional to the conditional mutual information I(token; context | category), and show this law holds across models, partitions, and over pretraining dynamics, including voiding a family of simplex ETF claims.