This is a stub page. It needs to be expanded with proper content.

Definition

ggml is a tensor library written in C/C++ designed for efficient machine learning inference on consumer hardware. It is the foundational compute backend for llama-cpp and derivative inference engines.

Key Points

  • Optimized for CPU and GPU inference with quantized weight formats
  • Enables local open-weight-models execution without cloud dependencies
  • Extended by community forks (e.g. strata-inference-engine) for MoE expert offload

Sources