AI Safety Index — Summer 2026 Edition
AI experts rate leading AI companies on key safety and security domains.
Overall Grades (July 2026)
| Company | Grade | Score |
|---|---|---|
| Anthropic | C+ | 2.66 |
| OpenAI | C | 2.28 |
| Google DeepMind | C | 2.01 |
| Meta | D+ | 1.32 |
| Z.ai | D- | 0.88 |
| Alibaba Cloud | D- | 0.87 |
| xAI | F | 0.65 |
| DeepSeek | F | 0.47 |
| Mistral | F | 0.33 |
Key Findings
Anthropic again earns the highest overall grade and leads five of six domains via relatively strong transparency, a comparatively established safety framework, technical research, and governance. OpenAI now leads in Risk Assessment on the strength of a broader evaluation suite and diverse engagement with external testing.
Meta improved from 6th to 4th place, while xAI dropped from 4th to 7th place.
Although the European Union is a leader in AI safety regulation, the top European AI company Mistral scored dead last on safety.
Three companies receive failing grades, one each from the US (xAI), China (DeepSeek), and Europe (Mistral).
Reviewers flagged the industry’s pivot to military AI use as an emerging current harm risk. From 2024 to 2026, companies including Anthropic, OpenAI, Google DeepMind, and Meta that previously banned military applications gradually reversed course.
Even industry leaders in safety practices are retreating from prior commitments. Anthropic, OpenAI, Google DeepMind, and Meta have weakened or voided pledges to pause unilaterally if redlines are approached, some citing competitor-contingent conditions. Reviewers call this “moving goalpost” behavior.
Existential Safety is the weakest domain industry-wide. No company exceeds C-; most score D or below.
Safety rhetoric outpaces revealed behavior. Across Google DeepMind, OpenAI, and xAI, leadership’s reassuring public messaging diverges from commercial conduct and legislative stance.
Companies are publishing and updating safety frameworks, but these frameworks have weak teeth — lacking quantitative thresholds, genuinely independent audits, and clear decision authority.
Methodology
The Summer 2026 Index evaluates nine leading AI companies on 37 indicators spanning six critical domains: Risk Assessment, Current Harms, Safety Frameworks, Existential Safety, Governance & Accountability, and Information Sharing.
An independent panel of seven leading AI researchers and governance experts reviewed company-specific evidence and assigned domain-level grades (A–F). The Index collected evidence up until June 3, 2026, combining publicly available materials with responses from a targeted company survey.
Prof. Stuart Russell (UC Berkeley): “While there is good work being done on AI safety in the industry, the capabilities race has become more extreme. Companies have backed away from earlier commitments to release new systems only with safety measures appropriate for their capability levels.”
Prof. David Krueger (University of Montreal): “AI companies’ lack of progress towards credible AI Safety plans is scandalous.”