Definition

Inference technique for large mixture-of-experts models where expert weight matrices reside primarily in system RAM and only frequently activated experts are cached in GPU VRAM — enabling models far larger than VRAM capacity on consumer hardware.

Key Points

  • 2026-10-04: strata-inference-engine uses expert offload (-cmoe class flags) for qwen3-8-flash-next — 512 experts, subset active per token (2026-10-04-strata-ferstar-4090-benchmark)
  • RTX 4090 24GB + 125GB RAM: VRAM holds hot experts + attention; cold experts in pinned memory
  • Decode speed depends on host RAM bandwidth and quant level (IQ2_XS, IQ3_S) as much as GPU tier
  • KV cache offload (--kv int8 --kv-resident) extends context beyond VRAM at latency cost

Sources