toxgate-cbb0a19b·1 events·first seen Aliases: ToxGate
A new arXiv preprint introduces ToxGate, a trust-fusion head that conditions external toxicity signals (English toxicity scores, Indic abuse priors, rule-based severity cues) on encoder representations before incorporating them into abuse detection predictions. The method is evaluated across three short-text abuse datasets, four transformer encoders, and five seeds, improving over baseline encoders in 10 of 12 in-domain and 7 of 8 transfer settings. The core finding is that external toxicity tools should be treated as conditional evidence rather than fixed features, with the largest gains in high-risk slices such as explicit slurs, violent threats, and cross-dataset transfer. The work targets a practical gap in content moderation for Indian multilingual and code-mixed text.