API model ID: deepseek-flash. 552B MoE for long tool-heavy agent workloads; up to 1M context.
CED separates encoder/decoder; CSA2 reuses sparse-attention state; FP4 caching targets agent cost shape: large prompts, repeated cache reads, smaller generated actions.
890 bytes/token global KV (~1/4 V4 Flash). Long-run quality depends on prompt structure, tool feedback, compaction, reasoning effort — context window is capacity limit, not quality guarantee.