modus-decoder-only-any-to-any-modeling-of-diverse-modalities-ac71ba9e·1 events·first seen Aliases: MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
Researchers at EPFL introduce Modus, a decoder-only model that treats all modalities symmetrically for any-to-any prediction without modality-specific heads, losses, or task pipelines. Unlike prior any-to-any models that use encoder-decoder or diffusion architectures trained from scratch, Modus leverages pre-trained decoder-only models as a prior. The model supports chained generation through intermediate modalities and cross-modal self-verification, achieving competitive performance with specialist and multitask baselines across benchmarks. All materials are open-sourced.