wordpiece-f55dc9e3·2 events·first seen Aliases: WordPiece
Researchers identify that English-centric byte-level tokenizers cause catastrophic autoregressive collapse when lightweight ASR models like Moonshine are applied to morphologically rich languages like Bengali. They propose a vocabulary transplantation pipeline that swaps in BanglaBERT's WordPiece vocabulary, reducing token fertility from 9.16 to 1.30 and cutting autoregressive sequence length by 85.8%. The modified model achieves 21.54% WER and an RTF of 0.0053 on the 882-hour Lipi-Ghor dataset, offering a reproducible blueprint for cross-script adaptation of compact ASR models without retraining from scratch.
ToaST (Tokenization with Split Trees) is a new subword tokenization method that uses a recursive binary split-tree inference procedure and Integer Programming-based vocabulary selection to directly optimize compression. On English text, ToaST reduces token counts by more than 11% compared to BPE, WordPiece, and UnigramLM at vocabulary sizes of 40,960 and above, effectively extending context length for models using it. In 1.5B parameter LM training experiments, ToaST achieves the highest CORE benchmark score, outperforming baselines by 2.6%–7.6% across 22 tasks. The LP relaxation of the vocabulary selection IP is near-integral in practice, yielding provably near-optimal vocabularies.