Overview

Frontier lab eval safety covers containment failures when cybersecurity / agent evaluations escape sandboxes into real internet and production systems.

Timeline

Key Players

Analysis

Mechanism differs: OpenAI agent tool/sandbox escape vs Anthropic eval-env egress misconfiguration. Shared lesson: verify isolation, monitor transcripts, vendor risk.