Overview
llama.cpp is an open-source C/C++ inference engine for open-weight-models, built on the ggml tensor library. Foundation for many local inference tools including ollama and engines like strata-inference-engine.
Recent Developments
- 2026-10-04: strata-inference-engine builds on llama.cpp/ggml with MoE expert offload extensions for qwen3-8-flash-next (2026-10-04-strata-codekk-overview)
- Community forks add
-cmoe, MTP, custom quant pipelines for frontier MoE models - Standard llama.cpp without offload cannot run 125B-class MoE on 12–24GB VRAM