OpenAI Releases LifeSciBench: Measuring AI Systems’ Capabilities in Real-World Scientific Research Scenarios
OpenAI has released LifeSciBench, a new evaluation benchmark designed to assess AI systems’ capabilities in real-world scientific research scenarios. LifeSciBench comprises 750 expert-crafted tasks covering seven categories of scientific workflows and seven biology domains. The tasks were contributed by 173 researchers holding doctoral degrees and possessing industry experience in biotechnology or pharmaceuticals. This benchmark emphasizes the assessment of complex scientific research competencies—including evidence integration, experimental design, data analysis, scientific reasoning, and scientific communication—rather than isolated factual questions. Over 79% of the tasks require multi-step reasoning, with an average of approximately four reasoning steps per task, and include 1,062 authentic research-related data attachments (e.g., papers, figures, sequence data, and structural files).