learning-to-hear-hesitation-continual-learning-for-disfluency-aware-asr-3cd7bf2f·1 events·first seen Aliases: Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR
A new arXiv preprint addresses the challenge of transcribing disfluent speech (hesitations, repetitions, fillers) in ASR systems, which typically omit such markers causing information loss. The authors introduce explicit disfluency tokens into a pretrained ASR model and apply continual learning to adapt across datasets with varying disfluency distributions while mitigating catastrophic forgetting. The work identifies a trade-off between disfluency marker learning and general ASR performance, and finds a consistent cross-attention head mechanism shared across continual learning methods.