Overview
Strata (Niko1221/Strata) is an open-source local inference engine running qwen3-8-flash-next on consumer GPUs (12GB+ VRAM) with 32–64GB+ system RAM. MIT licensed; builds on llama-cpp/ggml components.
Recent Developments
- 2026-10-04: HN visibility; claims 94 tok/s on RTX 5070 (project README); independent 4090 benchmark ~100–110 tok/s decode with IQ3_S (2026-10-04-strata-ferstar-4090-benchmark)
- OpenAI/Anthropic-compatible API at
http://127.0.0.1:8080/v1; MCP agent integration - One-click installer for Windows/Linux; quant selection by RAM tier (IQ2_XS, IQ3_S, Coder variant)
- Model compression: ISTA-DASLab, UkisAI Swift 1.5, Unsloth
Warning
Benchmark claims are project-published or single-user reports. Aggressive quantization (IQ2/IQ3) trades quality for VRAM fit. “125B on 12GB” requires expert offload to system RAM — not full-precision inference.