Definition
Inference technique for large mixture-of-experts models where expert weight matrices reside primarily in system RAM and only frequently activated experts are cached in GPU VRAM — enabling models far larger than VRAM capacity on consumer hardware.
Key Points
- 2026-10-04: strata-inference-engine uses expert offload (
-cmoeclass flags) for qwen3-8-flash-next — 512 experts, subset active per token (2026-10-04-strata-ferstar-4090-benchmark) - RTX 4090 24GB + 125GB RAM: VRAM holds hot experts + attention; cold experts in pinned memory
- Decode speed depends on host RAM bandwidth and quant level (IQ2_XS, IQ3_S) as much as GPU tier
- KV cache offload (
--kv int8 --kv-resident) extends context beyond VRAM at latency cost