shieldgemma-a9f7bdea·2 events·first seen Aliases: ShieldGemma
A new arXiv preprint evaluates whether fine-tuned encoder classifiers from the ModernBERT family (ModernBERT and Ettin) can replace LLM-based safety judges for detecting harmful outputs in user-model conversations. The study benchmarks encoders against rule-based methods, fine-tuned LLM classifiers, and LLM judges including LlamaGuard 3/4, ShieldGemma, StrongReject, and Claude-as-a-judge across multiple adversarial attack types. Results are reported on F1, false negative rate, and precision-recall, with breakdowns by attack technique, providing practical guidance on cost-latency tradeoffs for production safety pipelines.
Google released three new additions to the Gemma ecosystem: Gemma 2 2B, a small open-weights language model; ShieldGemma, a safety-focused classifier model; and Gemma Scope, an interpretability toolset. These releases expand the Gemma family with a smaller, more accessible model alongside dedicated safety and interpretability infrastructure. The announcement was published on the Hugging Face blog, indicating integration with the HF ecosystem.