Definition
Architectural techniques claimed by moonshot-ai for kimi-k3: Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), presented as enabling ~2.5× better scaling efficiency versus the K2 generation at 2.8T MoE scale with 1M-token context.
Key Points
- Bundled with Stable LatentMoE-style sparse expert routing (16 of 896 experts activated per MarkTechPost)
- Positioned as long-context and multimodal (native vision) enablers for frontier open models
- Full technical report and independent replication pending weight release (Jul 27, 2026)
Efficiency multipliers are vendor-claimed until open weights and third-party ablations exist.
Related
- kimi-k3
- mixture-of-experts
- attention-mechanism
- transformer-architecture
- open-weight-models
- moonshot-ai