June 30, 2026 — OpenAI released GeneBench-Pro, a benchmark designed to test whether AI models can make the judgment calls required in computational biology research.
The benchmark includes 129 problems across genomics, quantitative biology, and translational medicine. Each problem provides models with a dataset, experimental context, and a target question requiring data exploration, analytical approach selection, and iterative experimentation.
OpenAI sent 82 of 129 questions to external domain experts for realism review. Problems are built synthetically with controlled data-generation enabling grading against known targets.
Benchmark scores (OpenAI-reported):
- GPT-5.6 Sol: 28.7% pass rate (highest reasoning); 31.5% with Pro mode
- GPT-5: below 5% on original GeneBench
- Opus 4.8: 16.0%; Gemini 3.5 Flash: 8.1%; Gemini 3.1 Pro: 3.1%
- Grok 4.3: 1.5%; GLM 5.2: 4.6%; DeepSeek V4 Pro: 2.4%
Human experts estimate 20–40 hours per problem (~8,000 labor cost at $200/hour). Inference costs are several dollars per problem.
OpenAI is open-sourcing 10 representative questions on Hugging Face and providing a 50-question subset to Artificial Analysis for independent benchmarking.