Independent benchmark blog post (2026) documenting Qwen3.8-Flash-Next deployment via Strata on RTX 4090 D 24GB with 125 GiB DDR5 RAM (i9-14900KF, Z790).

Hardware context: 24GB VRAM limits expert caching to frequently used experts in GPU memory; full expert set stored in system RAM. Qwen3.8-Flash-Next is a 512-expert MoE model; BF16 weights exceed 300GB but only a subset activates per token.

Reported performance with IQ3_S quantization:

  • Decode: ~100–110 tok/s median
  • Cold prefill on 20–30k token prompts: 2500–2800 tok/s
  • Context: 262144 tokens (training length) with --kv int8 --kv-resident 32768 — QSA KV cells partially resident in pinned memory (~3.09 GiB), VRAM stable ~23900 MiB

Configuration notes: setup default --max-context 131072; author extended to 262144 without YaRN. CPU lacks AVX-512 so Q2_0 fastest kernel unavailable on this hardware.

Comparison: author migrated from Ollama; Strata decode reportedly “much faster” for code-scanning workloads on this hardware, though models and quantizations differ.