esmfold2-a9a4d0e0·2 events·first seen Aliases: ESMFold2
Biohub and EvolutionaryScale released ESMFold2, a 6.2-billion-parameter open-weights model for predicting the 3D shapes of proteins, DNA, RNA, and small molecules by treating molecular sequences as language. Unlike AlphaFold 3, ESMFold2 can operate without multiple sequence alignments (MSAs) by using a transformer-based embedding model (ESMC) trained on 2.8 billion sequences, outperforming Chai-1 in MSA-free settings and matching AlphaFold 3 when MSAs are provided. The model weights are freely available on HuggingFace and via API through Biohub, making frontier-level structural biology accessible without proprietary infrastructure. The release is significant for drug discovery involving novel or synthetic molecules where MSA databases may be sparse.
A Latent Space interview/commentary piece featuring Alex Rives of BioHub discussing ESMFold2 and the application of the 'bitter lesson' (scale and general methods beating hand-crafted inductive bias) to protein structure prediction and biology. The piece covers the tension between dataset scale versus domain-specific inductive bias in biological ML, and touches on world models and programmable biology. This represents a significant perspective from a leading researcher in protein language models on the next generation of biological foundation models.