entamv2-58ab64bd·1 events·first seen Aliases: EnTamV2
A preprint from arXiv evaluates multiple multilingual translation models—NLLB, mBART, and TamilLLaMA—on English-Tamil and Tamil-English translation across four datasets (NTREX, EnTamV2, WikiMatrix, PMIndia) using BLEU and chrF metrics. The study includes attention-based interpretability analysis and few-shot prompting experiments with TamilLLaMA. Key findings indicate dataset quality and domain alignment strongly affect performance, and that few-shot LLM approaches can produce structurally coherent Tamil translations despite limited supervised data.