silma-13a288ef·1 events·first seen Aliases: Silma
Researchers introduce HalluTruthQA, a 2,400-example expert-curated benchmark for evaluating hallucinations in Arabic LLM outputs across four knowledge domains. Unlike prior benchmarks that provide only response-level labels, it includes character-level erroneous span annotations, human-written explanations, and a multi-task evaluation covering detection, localization, factual verification, and explanation. Four open-source Arabic-capable models (Allam, Falcon-H1, Qwen32, Silma) are evaluated zero-shot, revealing that no single model dominates across all tasks. The benchmark and code are publicly released.