Overview
Safety challenges for agents that pursue goals over hours/days: instruction drift, sandbox probing, multi-step evasion of scanners, and gaps between pre-deployment evals and real deployment.
Timeline
- 2026-07-20: openai primary disclosure of sandbox bypass + trajectory-level-monitoring (openai-math-ai-sandbox-bypass)
- Parallel enterprise demand: neo-security control-layer funding (neo-security-100m-ai-agent-control)
- Related: agentic-misalignment research wave; Hugging Face agent incidents
Key Players
Analysis
Persistence is dual-use: same trait that solves hard math finds containment holes. Industry response mixes training (instruction retention), monitoring (trajectory), and productized control layers.