OpenAI has released GeneBench-Pro, a new benchmark designed to evaluate AI performance on genomics, biology, and scientific research tasks using complex, real-world datasets. The announcement comes from OpenAI's official blog, positioning it as a domain-specific evaluation tool for scientific AI. This extends the evaluation landscape into life sciences, a growing area of AI application.
OpenAI has released LifeSciBench, a benchmark designed to evaluate AI systems on real-world life science research tasks and decisions. The benchmark is described as expert-authored and expert-reviewed, targeting domain-specific evaluation in biology and related fields. This addresses a gap in specialized scientific benchmarking for AI systems.
OpenAI has released HealthBench, a new evaluation benchmark designed to assess AI model performance and safety in healthcare settings. The benchmark was developed with input from over 250 physicians and targets realistic clinical scenarios. It aims to establish a shared standard for measuring how well AI models handle health-related tasks.
OpenAI introduces a real-world evaluation framework designed to measure how AI systems can accelerate biological research in wet lab settings. The work uses GPT-5 to optimize a molecular cloning protocol as a concrete demonstration case. The framework explicitly addresses both the potential benefits and biosecurity risks of AI-assisted experimentation, positioning this as a dual-use capability assessment.
OpenAI introduces PaperBench, a benchmark designed to evaluate AI agents' ability to replicate state-of-the-art AI research papers end-to-end. The benchmark targets a high-complexity capability: reproducing experimental results from frontier AI research, which requires code generation, experimental design, and scientific reasoning. This positions PaperBench as a tool for tracking progress toward autonomous AI research agents.
OmniaBench is a new benchmark for evaluating general AI agents across diverse scenarios, spanning 90 level-1 and 354 level-2 domains derived from app stores, product documents, and web retrieval. The benchmark contains 1,431 tasks with single-turn and multi-turn formats, a ten-dimensional capability taxonomy, and eight atomic difficulty factors for fine-grained analysis. Even frontier models like Claude Sonnet 5 and GPT-5.6-Sol achieve only ~58% Overall Pass@1, revealing persistent weaknesses in planning, constraint maintenance, and adaptive correction. The work addresses a gap in existing agent benchmarks that tend to cover narrow tool ecosystems or interaction formats.
OpenAI has released FrontierScience, a new benchmark designed to evaluate AI reasoning capabilities across physics, chemistry, and biology. The benchmark is intended to measure progress toward AI systems capable of performing real scientific research tasks. This represents OpenAI's effort to establish a rigorous evaluation framework for frontier-level scientific reasoning, going beyond standard academic problem sets.
OpenAI released Procgen Benchmark, a suite of 16 procedurally-generated environments designed to measure how quickly reinforcement learning agents learn generalizable skills. The benchmark targets a core challenge in RL: distinguishing memorization of specific environments from genuine skill generalization. Its procedural generation ensures agents cannot overfit to fixed level layouts.
OpenAI published an analysis identifying problems with SWE-Bench Pro, a widely-used benchmark for evaluating AI coding capabilities. The post raises concerns about the benchmark's reliability and accuracy as a signal for model performance. This matters because SWE-Bench is a primary reference point for comparing frontier coding models, and benchmark integrity directly affects how labs and practitioners interpret capability claims.