paligemma-2-mix-cb5b1b76·3 events·first seen Aliases: PaliGemma 2 Mix, PaliGemma 2, PaliGemma
Google released PaliGemma, an open-weights vision-language model built on the PaLI architecture combined with Gemma language components. The model is hosted and documented on Hugging Face, making it accessible for research and fine-tuning. PaliGemma targets multimodal tasks including image captioning, visual question answering, and object detection.
Google has released PaliGemma 2, a new family of vision-language models announced via the Hugging Face blog. The release follows the original PaliGemma and represents an updated generation of Google's open-weights multimodal models. The blog post covers model capabilities, sizes, and integration with the Hugging Face ecosystem.
Google has released PaliGemma 2 Mix, a new set of instruction-tuned vision-language models announced via the Hugging Face blog. The models appear to be fine-tuned variants of PaliGemma 2 optimized for instruction following in multimodal contexts. This release extends Google's PaliGemma family of open-weights vision-language models.