mmar-7b889108·2 events·first seen Aliases: MMAR
X³-OPD is a new training framework that distills reasoning capabilities from a powerful text-based teacher model into an audio-language student model via on-policy alignment. The approach generates reasoning trajectories conditioned on the student's own acoustic perception while the teacher provides token-level guidance from matched textual inputs. A three-tier symmetric corpus covers speech-rendered text reasoning, audio-event reasoning, and paralinguistic spoken-dialogue reasoning. Evaluations on MMSU, MMAU, BIG Bench Audio, and MMAR show substantial improvements in audio-grounded reasoning and chain-of-thought quality.
Researchers introduce AudioDER, a ~191k-sample post-training dataset for Large Audio-Language Models (LALMs) built via an acoustic similarity-based deduplication pipeline to reduce redundancy and improve corpus diversity. Each sample pairs an audio clip with a multiple-choice question, answer candidates, a caption, and a chain-of-thought rationale generated by Qwen3-30B. Post-training Qwen2-Audio-7B-Instruct on AudioDER yields consistent gains on audio reasoning benchmarks including MMAU-mini, MMSU, and MMAR. The work addresses a data quality gap in audio-language training rather than proposing a new model architecture.