Overview

Qwen3.8-Flash-Next is a 125B-parameter mixture-of-experts model from the qwen family (alibaba). Uses 512 experts with sparse activation per token; BF16 weights exceed 300GB uncompressed.

Recent Developments

  • 2026-10-04: Runnable on consumer hardware via strata-inference-engine with aggressive quantization and expert RAM offload (2026-10-04-qwen-3-8-flash-strata-consumer-hardware)
  • GSQ-RCO quantization (ISTA-DASLab); community GGUF variants from Unsloth
  • Training context length: 262144 tokens
  • Commercial licensing terms apply — verify Alibaba model license for production use

Sources