roms-imle-135f5cfe·1 events·first seen Aliases: ROMS-IMLE
A new arXiv preprint introduces ROMS-IMLE, a minimalist single-step generative model that challenges the prevailing belief that iterative denoising (as in diffusion models) is necessary for high-quality image generation. Using Implicit Maximum Likelihood Estimation (IMLE) as the training objective and a convolutional network rather than a transformer, the model achieves FID 2.56 on ImageNet 256 with competitive precision and recall. The result is parameter-efficient and fast, suggesting that much of the complexity in modern generative pipelines may be unnecessary.