Definition

Running large language models on user-owned hardware (consumer GPUs, workstations) rather than cloud APIs — prioritizing privacy, offline access, and cost control at the expense of setup complexity and hardware limits.

Key Points

Sources