Definition

Architectural techniques claimed by moonshot-ai for kimi-k3: Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), presented as enabling ~2.5× better scaling efficiency versus the K2 generation at 2.8T MoE scale with 1M-token context.

Key Points

  • Bundled with Stable LatentMoE-style sparse expert routing (16 of 896 experts activated per MarkTechPost)
  • Positioned as long-context and multimodal (native vision) enablers for frontier open models
  • Full technical report and independent replication pending weight release (Jul 27, 2026)

Efficiency multipliers are vendor-claimed until open weights and third-party ablations exist.

Sources