
arxiv-19dd39d0·7 events·first seen Aliases: ArXiv
Researchers introduce SciDiagramEdit, a benchmark and skill-evolution framework for automating edits to scientific figures under natural-language instructions. The benchmark mines before/after figure pairs from arXiv version histories, grounding training in authors' own revision intent. An agentic proposer iteratively refines the agent's skill specification from execution traces, progressively improving edit accuracy on held-out validation. The work demonstrates that natural paper revisions are an effective training signal for instruction-driven figure editing.
Researchers present a multi-phase LLM pipeline that autonomously processes 11,083 condensed-matter physics arXiv papers, conceives a research direction, calibrates methodology against published references, runs first-principles computations, and produces a publication-grade manuscript on altermagnetic piezomagnetism. The system operates across 47 fresh-context sessions with 2,162 literature-consultation events, using redundancy and adversarial review for fault tolerance, with human intervention limited to operational knowledge curation at reproduction failures rather than scientific direction. Ablation experiments isolate 'structurally enforced numerical confrontation' at calibration checkpoints as the key mechanism preventing hallucination in physical science domains. The work extends autonomous research agents beyond ML sandboxes into high-stakes physical science, where external literature anchors replace execution-based calibration.
COMPOSE is a framework that generates plausible future mathematical theorem-like claims by conditioning a language model on both a scientific citation graph and a formal theorem dependency graph simultaneously. The authors construct a dataset of 108K paired scientific-formal graph examples from arXiv and Mathlib, plus a benchmark of 47K future papers from 2024–2025. Experiments show COMPOSE outperforms baselines on retrieval to real future papers and LLM-judge evaluation, producing more grounded and mathematically richer outputs. The work advances AI-assisted mathematical reasoning by combining informal scientific context with formal proof structure.
A preprint on arXiv proposes a sleep-like memory consolidation mechanism for large language models, drawing an analogy to biological sleep-based memory consolidation in neural systems. The work appears to address how LLMs might better retain and integrate new information over time, a key challenge in continual learning and knowledge updating. The paper attracted notable community attention on Hacker News with 164 points and 122 comments, suggesting broad interest in the approach.
AiraXiv is a proposed open-access academic publishing platform designed to accommodate both human and AI-generated research outputs, addressing scalability challenges in traditional peer review. The platform supports AI scientists via Model Context Protocol (MCP)-based interactions and human scientists through an interactive UI, with papers evolving through continuous feedback-driven iteration. It was validated through real-world deployment as the submission platform for ICAIS 2025. The work positions itself as infrastructure for a future where AI agents are first-class participants in the scientific publishing ecosystem.
Hugging Face announced an integration allowing ML demos to be linked or embedded directly on arXiv paper pages. This lowers the barrier between research publication and interactive model demonstration. The feature connects academic papers to live Spaces or model demos hosted on Hugging Face.
A Ramp AI Index survey shows Anthropic reached 34.4% business adoption in April 2026, surpassing OpenAI's 32.3%, though analysts cite token cost inflation, service degradation, and competition from cheaper inference platforms as threats to the lead. Cerebras surged 89% on its IPO debut, signaling investor appetite for AI infrastructure hardware. Separately, Anthropic's withheld Claude Mythos model—which solved a novel cybersecurity challenge—prompted meetings with the Financial Stability Board, while ArXiv announced year-long bans for authors submitting unvetted AI-generated content.