This page may contain stale information. Last updated: 2026-08-14
Overview
Large-scale cybersecurity evaluation suite for AI agents on real-world vulnerability analysis (1,500+ instances / 188 projects per Microsoft model card). Also linked to ExploitGym tasks in the openai / hugging-face rogue-agent incident narrative.
Recent Developments
- 2026-08-14: zhipu-ai reports glm-5-3 at 84.5% CyberGym (vs glm-5-2 77.2%); company comparison vs Mythos 5 / GPT-5.6 Sol (2026-08-14-zai-glm-53-unite-ai)
GLM CyberGym scores are vendor-reported pending independent replication.