shieldstral-46c47cb3·1 events·first seen Aliases: Shieldstral
Shieldstral is a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7× its size on text safety benchmarks and claims state-of-the-art on multimodal safety classification. The model frames content moderation as a binary yes/no question-answering task, allowing heterogeneous safety datasets with divergent taxonomies to be unified under a single training framework. The authors describe a data construction pipeline covering ~54.1M samples and a fine-grained evaluation set for policy adaptability. The efficiency-at-scale result is notable for practitioners deploying content moderation in resource-constrained settings.