automated-reproducibility-assessments-in-the-social-and-behavioral-sciences-using-large-language-models-6323e3a8·1 events·first seen Aliases: Automated reproducibility assessments in the social and behavioral sciences using large language models
A preprint from arXiv demonstrates that an LLM pipeline can automate reproducibility assessments of published social and behavioral science studies, recovering original effect sizes in 41% of cases (vs. 34% for human reanalysts) and reaching the same qualitative conclusion in 96% of cases (vs. 74% for humans). The study evaluated 76 published studies with predefined claims. The results suggest LLMs could serve as a scalable tool for systematic auditing of empirical research, addressing the resource-intensive nature of traditional reproducibility efforts.