A new arXiv preprint evaluates how well LLMs (ChatGPT, Claude, DeepSeek) can generate one-page research project plans in physics, astrophysics, and cosmology, comparing them against human-written proposals across 32 total documents. Human reviewers rated human and AI proposals similarly and correctly identified authorship ~72-79% of the time, while AI reviewers (Claude Opus 4.8 and ChatGPT Pro 5.5) correctly classified all 32 proposals and scored AI-written proposals roughly one point higher than human-written ones on a five-point scale. The study raises a concrete concern about deploying LLMs in grant proposal evaluation pipelines, as AI reviewers exhibit a systematic preference for AI-generated content.
A controlled study evaluated eight expert-defined research projects in physics, astrophysics, and cosmology, comparing literature reviews performed by human experts against those by ChatGPT-4o, ChatGPT Deep Research, and Gemini. Human-AI reference overlap was below 6%, and 64% of AI-generated references had metadata errors (incorrect title, author, year, etc.), though only 3% were fully fabricated. A preliminary test of GPT-5.5 showed zero fabrications or metadata mismatches, suggesting significant improvement in the 2026 generation. The findings indicate mid-2025 models are complementary rather than substitutes for expert literature search, and require systematic verification.
SoundnessBench is a new benchmark of 1,099 machine-learning research proposals derived from ICLR submissions, labeled with reviewer soundness scores, designed to test whether LLMs can reliably distinguish methodologically sound research ideas from unsound ones. Evaluated across 12 frontier LLMs, the benchmark reveals a pervasive optimism bias: models systematically rate low-soundness proposals as sound under standard prompting, with aggressive prompting shifting errors from false positives to false negatives rather than eliminating them. Controls for data contamination, surface features, and human audit quality suggest the bias is not attributable to a single confounder. The authors conclude that current LLMs are not yet reliable as standalone first-gate evaluators of scientific rigor, a critical bottleneck for autonomous AI research agents.
A new arXiv paper introduces a large-scale evaluation framework for comparing LLM-generated research ideas against human-authored ones, using reverse-engineered prior-work sets as prompts. The authors develop a two-axis taxonomy of research taste (opportunity pattern and research paradigm) and find a consistent distributional gap: LLMs over-index on bridge-like opportunities and synthesis methods, while human researchers spread more broadly across framing and contribution types. The result suggests current LLMs produce reasonable but systematically narrower and shifted ideation relative to human researchers.
A new arXiv preprint tests the implicit assumption that LLM evaluation is easier than generation, using a controlled in-context QA setup across four benchmarks (SQuAD 2.0, DROP, HotpotQA, MuSiQue) and two models. Results show generation accuracy exceeds self-evaluation accuracy on three of four benchmarks, with attention analysis revealing that evaluation attends to context 3–5x less than generation does. LoRA fine-tuning experiments confirm the asymmetry is not a training artifact, with cross-task interference observed in both directions. The findings directly challenge assumptions underlying LLM-as-a-Judge and self-evaluation pipelines widely used in RLHF and agentic systems.
A preprint from arXiv demonstrates that an LLM pipeline can automate reproducibility assessments of published social and behavioral science studies, recovering original effect sizes in 41% of cases (vs. 34% for human reanalysts) and reaching the same qualitative conclusion in 96% of cases (vs. 74% for humans). The study evaluated 76 published studies with predefined claims. The results suggest LLMs could serve as a scalable tool for systematic auditing of empirical research, addressing the resource-intensive nature of traditional reproducibility efforts.
A new arXiv paper argues that standard LLM benchmarks overstate model capabilities by focusing on average performance on training-data-adjacent tasks while ignoring response variance and error magnitude. The authors introduce a novel benchmark requiring frontier LLMs to write code for data analysis tasks, comparing results against human expert submissions. Human experts outperformed the frontier LLM on average across multiple metrics and showed lower performance variability. The findings challenge the prevailing narrative that LLMs perform at human-expert level on knowledge economy tasks.
A preprint uses an open problem from EC 2025 as a testbed to evaluate AI-assisted research workflows in economics and computer science. The study examines whether human intuition in prompts, multi-turn interaction, and LLM capability compare favorably to a first-year PhD student's contributions. Key findings: human intuition in prompts improves LLM 'taste', multi-turn workflows help when encouraging ambitious steps, and the LLM performs slightly below the first-year PhD student on the same problem. The work contributes empirical evidence on the practical utility and limits of LLMs as research collaborators in formal theory domains.
New research suggests that large language models not only inherit human biases from training data but can also develop novel biases of their own when used in hiring contexts. The study raises concerns about AI résumé screening systems that operate before any human review. This adds to a growing body of evidence that LLM-based hiring tools may produce unfair outcomes in ways that are difficult to anticipate or audit.