Definition

A zero-shot attack that injects fabricated chain-of-thought text styled like a model’s internal reasoning so the model treats the forgery as its own thoughts and complies with harmful instructions.

Key Points

  • ~60% attack success vs near-zero baseline across frontier models (2026-07-30-cot-forgery-arxiv)
  • Won OpenAI Aug 2025 red-teaming hackathon; similar to OpenAI “fake chain of thought” / gpt-red findings
  • Demonstrates role-confusion: style outweighs role tags
  • Dual-use: describe mechanism conceptually; avoid attack recipes beyond published summaries

Sources