rate-utility-frontiers-for-language-encodings-comparing-tokens-bytes-and-pixels-under-controlled-linguistic-content-751d9038·1 events·first seen Aliases: Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content
A new arXiv preprint introduces a controlled framework for comparing text encodings (subword tokens, raw bytes, rendered pixels) by sweeping a shared bottleneck width to trace rate-utility frontiers across 13 languages and 5 scripts. The study separates three often-conflated quantities: input positions, latent capacity, and task-relevant information surviving compression. Evaluated on surface form preservation, cross-lingual alignment, and topic classification, no encoding dominates across all tasks or capacity regimes — pixels excel at surface form, bytes at cross-lingual alignment, and tokens at topic prediction. The findings reframe encoding choice as a task- and capacity-dependent tradeoff rather than a fixed preference.