content-is-what-remains-invariant-speech-tokenization-from-parallel-utterances-8cdb4ee5·1 events·first seen Aliases: Content is What Remains: Invariant Speech Tokenization from Parallel Utterances
Researchers introduce PINT (Parallel INvariant Tokenization), a method that fine-tunes a self-supervised speech encoder using alignment losses across parallel utterances from multiple speakers to extract purely linguistic content tokens. The approach achieves a 98.7% relative reduction in speaker probe accuracy and 42% lower ABX error rate compared to baselines, while preserving frame-level temporal grounding. PINT tokens are designed as drop-in semantic targets for audio codecs, addressing a known limitation of HuBERT-style SSL models that leak speaker identity and prosody into discrete representations.