Overview
On September 29, 2026, anthropic published research finding zhipu-ai’s open-weight glm-5-3 matches claude-mythos Preview on end-to-end exploit development but lacks robust safeguards. caisi (NIST) endorsed GLM-5.3 as the most cyber-capable open-weight model released to date.
Key Facts
- Safeguard bypass rates: 64–100% via deceptive prompts, thinking-token prefilling, or abliteration (~$4,400 cost)
- CAISI aggregate: GLM-5.3 lags US frontier by ~4 months on cyber benchmarks
- Anthropic urges expanded defender access and government safety testing
- Z.ai’s own blog acknowledges emergent cyber capability from scaled post-training
Related
- glm-5-3
- anthropic
- zhipu-ai
- ai-agent-security
- open-weight-models
- cyber-capability-evaluation-risk
- us-voluntary-ai-cyber-testing