donut-71f341aa·1 events·first seen Aliases: Donut
Researchers introduce Persian Pixel, a synthetic OCR dataset of over 343,000 image-text pairs covering sentence, paragraph, and full-page layouts generated from a 7-million-word Persian corpus. The dataset addresses the scarcity of annotated Persian OCR data by modeling typographic complexities of Perso-Arabic script and applying 25+ stochastic degradation models to simulate real-world document artifacts. It is designed to train and fine-tune transformer-based OCR architectures such as TrOCR and Donut, and is released openly to support Persian document digitization research.