Definition
Running large language models on user-owned hardware (consumer GPUs, workstations) rather than cloud APIs — prioritizing privacy, offline access, and cost control at the expense of setup complexity and hardware limits.
Key Points
- 2026-10-04: strata-inference-engine enables qwen3-8-flash-next (125B MoE) on 12GB+ VRAM cards via expert offload + aggressive quantization (2026-10-04-qwen-3-8-flash-strata-consumer-hardware)
- Ecosystem: ollama, llama-cpp, llama.cpp forks, LM Studio, vLLM (server-class)
- OpenAI-compatible local APIs (
localhost:8080/v1) enable agent tool integration without cloud dependency - Turkish AI enthusiast community actively experiments with consumer GPU setups (RTX 4090/5070 class)
Related
- strata-inference-engine
- ollama
- llama-cpp
- open-weight-models
- ai-inference
- consumer-gpu-ai
- expert-offload-inference