wavlm-3efd181e·2 events·first seen Aliases: WavLM
Researchers propose a text-free framework for second-language (L2) speech assessment using Dynamic Time Warping (DTW) over self-supervised WavLM representations, covering phonetic accuracy, rhythm, and intonation in English and Japanese. The DTW-based approach comparing learner speech to native templates exceeds human agreement on holistic phonetic scoring, and a novel warping-path method approaches human-level rhythm assessment. Intonation scoring, combining DTW over prosodic residuals with pitch and intensity features, shows more modest results. The method requires no labeled L2 data, making it applicable in low-resource settings.
Researchers propose an audio-native explainability pipeline using Integrated Gradients on time-aligned self-supervised representations to localize decision evidence in deepfake speech detectors. Applied to three WavLM-based detectors (AASIST, CA-MHFA, SLS) on the ASVspoof 5 benchmark, the method reveals that despite similar performance, each detector relies on fundamentally different cues: environmental noise, phoneme artifacts, and word boundaries respectively. Findings are validated via causal masking experiments that confirm performance degrades when primary cues are removed. The work advances interpretability of audio deepfake detection, relevant to AI safety and media authenticity.