Overview

Strata (Niko1221/Strata) is an open-source local inference engine running qwen3-8-flash-next on consumer GPUs (12GB+ VRAM) with 32–64GB+ system RAM. MIT licensed; builds on llama-cpp/ggml components.

Recent Developments

  • 2026-10-04: HN visibility; claims 94 tok/s on RTX 5070 (project README); independent 4090 benchmark ~100–110 tok/s decode with IQ3_S (2026-10-04-strata-ferstar-4090-benchmark)
  • OpenAI/Anthropic-compatible API at http://127.0.0.1:8080/v1; MCP agent integration
  • One-click installer for Windows/Linux; quant selection by RAM tier (IQ2_XS, IQ3_S, Coder variant)
  • Model compression: ISTA-DASLab, UkisAI Swift 1.5, Unsloth

Warning

Benchmark claims are project-published or single-user reports. Aggressive quantization (IQ2/IQ3) trades quality for VRAM fit. “125B on 12GB” requires expert offload to system RAM — not full-precision inference.

Sources