scaling-native-multimodal-pre-training-from-scratch-d51a84fd·1 events·first seen Aliases: Scaling Native Multimodal Pre-Training From Scratch
A new arXiv preprint investigates compute-optimal scaling for transformer-based vision-language models trained natively on multimodal inputs from scratch, rather than via late-fusion of separately pre-trained components. The authors derive power-law relationships between compute budget, model size, and token count, finding that language and multimodal objectives exhibit distinct scaling behaviors and that data mixture composition strongly influences optimal resource allocation. They also report positive cross-modal transfer effects, including improved pure-text spatial reasoning. The work establishes an efficiency frontier for configuring model size, token count, and data mixture under fixed compute.