This page may contain stale information. Last updated: 2026-08-14

Overview

Large-scale cybersecurity evaluation suite for AI agents on real-world vulnerability analysis (1,500+ instances / 188 projects per Microsoft model card). Also linked to ExploitGym tasks in the openai / hugging-face rogue-agent incident narrative.

Recent Developments

GLM CyberGym scores are vendor-reported pending independent replication.

Sources