x-opd-1d383170·1 events·first seen Aliases: X³-OPD
X³-OPD is a new training framework that distills reasoning capabilities from a powerful text-based teacher model into an audio-language student model via on-policy alignment. The approach generates reasoning trajectories conditioned on the student's own acoustic perception while the teacher provides token-level guidance from matched textual inputs. A three-tier symmetric corpus covers speech-rendered text reasoning, audio-event reasoning, and paralinguistic spoken-dialogue reasoning. Evaluations on MMSU, MMAU, BIG Bench Audio, and MMAR show substantial improvements in audio-grounded reasoning and chain-of-thought quality.