qwen3-30b-57a1c433·2 events·first seen Aliases: Qwen3-30B
Researchers introduce AGC-Bench, a comprehensive AI creativity benchmark built from a systematic review of 3,101 papers and 497 existing benchmarks, covering 78 datasets across brainstorming, STEM, narrative, figurative language, and humor. The work introduces Judge Response Theory to correct for LLM-as-judge bias and fine-tunes Qwen3-30B to produce AGC-Judge, an open-weight scoring model. Key findings include the recovery of a single creativity factor 'c' (analogous to the general intelligence 'g' factor) explaining 81.5% of variance across 83 LLMs, and evidence that top humans still outperform top LLMs on creativity tasks. The benchmark, leaderboard, and human data are released as open infrastructure.
Researchers introduce AudioDER, a ~191k-sample post-training dataset for Large Audio-Language Models (LALMs) built via an acoustic similarity-based deduplication pipeline to reduce redundancy and improve corpus diversity. Each sample pairs an audio clip with a multiple-choice question, answer candidates, a caption, and a chain-of-thought rationale generated by Qwen3-30B. Post-training Qwen2-Audio-7B-Instruct on AudioDER yields consistent gains on audio reasoning benchmarks including MMAU-mini, MMSU, and MMAR. The work addresses a data quality gap in audio-language training rather than proposing a new model architecture.