Definition
Usage-based pricing for LLM APIs charged per input/output tokens. Unlike conventional software where buggy loops are cheap, inefficient prompts/retries can produce large, silent cost overruns.
Key Points
- 2026-07: amazon internal briefing cited $1.8M Claude Sonnet overrun (FT-attributed) prompting automated spend guardrails (2026-07-30-amazon-ai-overspend-pymnts)
- Related failure mode: gamified usage leaderboards (“tokenmaxxing”) on kiro
- Enterprise response: deployment-path caps, caching, model tiering (ai-spend-guardrails)