marbert-029cbaa0·1 events·first seen Aliases: MARBERT
A new arXiv preprint evaluates five Arabic language models (including MARBERT and AraBERT) against adversarial attacks at character, word, and sentence levels, finding that diacritic insertion can reduce model accuracy by up to 92% and paraphrasing attacks degrade performance by an average of 76%. The study also tests adversarial training as a defense, finding MARBERT most robust and AraBERT showing the greatest relative gains. Results highlight particular challenges for morphologically rich languages where character-level noise remains difficult to defend against.