This is a stub page. It needs to be expanded with proper content.
Definition
ggml is a tensor library written in C/C++ designed for efficient machine learning inference on consumer hardware. It is the foundational compute backend for llama-cpp and derivative inference engines.
Key Points
- Optimized for CPU and GPU inference with quantized weight formats
- Enables local open-weight-models execution without cloud dependencies
- Extended by community forks (e.g. strata-inference-engine) for MoE expert offload