trocr-9fb932a7·1 events·first seen Aliases: TrOCR
Researchers introduce Persian Pixel, a synthetic OCR dataset of over 343,000 image-text pairs covering sentence, paragraph, and full-page layouts generated from a 7-million-word Persian corpus. The dataset addresses the scarcity of annotated Persian OCR data by modeling typographic complexities of Perso-Arabic script and applying 25+ stochastic degradation models to simulate real-world document artifacts. It is designed to train and fine-tune transformer-based OCR architectures such as TrOCR and Donut, and is released openly to support Persian document digitization research.