Definition
A zero-shot attack that injects fabricated chain-of-thought text styled like a model’s internal reasoning so the model treats the forgery as its own thoughts and complies with harmful instructions.
Key Points
- ~60% attack success vs near-zero baseline across frontier models (2026-07-30-cot-forgery-arxiv)
- Won OpenAI Aug 2025 red-teaming hackathon; similar to OpenAI “fake chain of thought” / gpt-red findings
- Demonstrates role-confusion: style outweighs role tags
- Dual-use: describe mechanism conceptually; avoid attack recipes beyond published summaries