Moonshot AI (Beijing, Alibaba-backed) released Kimi K3 on Thursday, July 16, 2026 — a 2.8-trillion-parameter model it calls the world’s largest open-source AI model, performing near proprietary systems from Anthropic and OpenAI. Timing landed just ahead of the 2026 World Artificial Intelligence Conference in Shanghai.

Availability: API live now via kimi.com (Google account or phone; no credit card). Full model weights scheduled for July 27, 2026. OpenAI SDK-compatible API. Pricing: 15/M output; cached input 1,000 top-ups).

Architecture: Sparse MoE at 2.8T total parameters (~75% larger than DeepSeek V4 Pro ~1.6T). 1-million-token context window; native vision; always-on “thinking mode.” Built on Kimi Delta Attention (hybrid linear attention) and Attention Residuals (drop-in residual replacement). Both techniques previously published as open research on GitHub.

Benchmarks (company + Artificial Analysis / public boards):

  • GDPval-AA v2: 1,687 — 3rd overall behind Claude Fable 5 Max (1,815) and GPT-5.6 Sol Max (1,747.8); ahead of Claude Opus 4.8 (1,600)
  • AA-Briefcase: 1,527 — 2nd behind Fable 5 Max (1,587); ahead of GPT-5.6 Sol Max (1,495)
  • BrowseComp: 91.2/100 SOTA for long-horizon information seeking
  • Arena.AI Frontend Code Arena: #1 at 1,679 Elo
  • Led Automation Bench, SpreadsheetBench 2, BrowseComp among eight automation benchmarks; top-three across six coding benches; led SWE Marathon and Program Bench

Demos: 48-hour autonomous chip-design PoC (4 mm² design, 100 MHz timing convergence, >8,700 tokens/sec decode in simulation via open-source EDA). Astrophysics I-Love-Q relation reproduced in ~2 hours vs 1–2 weeks for a senior researcher.

Company context: Founded 2023 by Yang Zhilin. Raised ~2.5B→5B. User ranking fell after DeepSeek R1 (Jan 2025); open-source pivot via K2 (Jul 2025) and K2.5 (Jan 2026).

Product lineup: K3 flagship (15); K2.7 Code (4); K2.6 general (4). Kimi Code CLI (3,100+ GitHub stars) updated to 0.25.0/0.26.0 with subagents, background tasks, plan mode. VSCode/Cursor/Zed integration.

Corroboration (SCMP, Jul 17): Moonshot acknowledges overall performance still trails most powerful proprietary models but claims frontier-level results outperforming GPT-5.5, Claude Opus 4.8, GLM-5.2 on many evals; beat Claude Fable 5 and GPT-5.6 Sol on Program Bench and SWE Marathon (self-reported).