Hacker News discussion of Niko1221/Strata — open-source inference engine for Qwen3.8-Flash-Next on consumer GPUs.

Community benchmark reports (unverified, hardware-dependent):

  • RTX 4090 + 128GB DDR5 + Ryzen 7950X3D: ~124 tok/s reported by one user
  • RTX 3090 (power-capped, PCIe 8x, Zen 2, 64GB DDR4): 40–60 tok/s with optimization; code benchmark 114/128 vs 92/128 for ninfer-3090 variant

Skepticism in thread: widespread Strata promotion; some users report hype exceeding measured accuracy gains; quality vs speed tradeoffs with aggressive quantization (IQ2/IQ3 class).

Technical context from discussion: Strata uses expert offload (-cmoe style) parking experts in RAM; attention on GPU. Light quantizations (Q2_0, IQ2_XS) required for high decode speeds on 24GB cards. Comparison to Ollama and llama.cpp forks common.

Model: Qwen3.8-Flash-Next on Hugging Face; compressed by ISTA-DASLab, UkisAI, Unsloth; Strata builds on llama.cpp/ggml components.