self-supervision-drives-representational-convergence-in-medical-foundation-models-more-than-clinical-supervision-5748b130·1 events·first seen Aliases: Self-supervision drives representational convergence in medical foundation models more than clinical supervision
A controlled study across 18 image and 7 text encoders spanning 7M to 27B parameters finds that representational convergence in medical foundation models is primarily driven by the self-supervised pretraining objective, not by clinical supervision or model scale. Using 650,982 chest radiographs across six datasets, the authors show matched self-supervised encoders align at 40.4% versus 21.1% for label-supervised and 3.3% for image-text models, and that convergence does not grow with parameter count. Despite modest convergence, a linear classifier transfers across encoders to five held-out hospitals retaining ~85% of within-encoder performance, suggesting interoperability must be deliberately designed through objective choice rather than assumed from scale.