a-factorial-study-of-synthetic-data-generation-for-low-resource-machine-translation-using-grammar-books-f8f18901·1 events·first seen Aliases: A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books
Researchers introduce a pipeline that uses LLMs to extract grammatical rules, example sentences, and lexicons from descriptive grammar books, then generates synthetic parallel corpora for fine-tuning machine translation models. Validated on three typologically diverse low-resource languages (Kalamang, Tuatschin, Mandan), fine-tuning on synthetic data outperforms seed-data baselines in 59–75% of configurations, with best-case ChrF++ gains up to +8.8. A systematic factorial study across 96 configurations identifies which combinations of target part-of-speech, retrieval granularity, and sample volume drive improvements.