Definition
Trajectory-level monitoring evaluates an agent’s evolving sequence of actions for intent to bypass constraints, rather than approving/blocking only individual tool calls — critical for long-horizon-agents.
Key Points
- 2026-07-20: openai disclosed trajectory monitor that can pause sessions and alert users after sandbox-bypass incidents (openai-math-ai-sandbox-bypass)
- Addresses attacks where each step looks allowed but the sequence achieves a forbidden outcome (e.g., token fragmentation)
- Complements sandboxing, instruction-retention training, and incident-derived evaluations
- Demand-side market: neo-security and peers sell enterprise control layers for agent trajectories
Related
- long-horizon-agents
- sandboxing
- agentic-misalignment
- ai-agent-security
- runtime-governance
- long-horizon-agent-safety