Definition

Techniques that reduce model size, memory, and compute cost — including quantization, pruning, distillation, and tensor-networks — to enable edge-ai, on-prem, and cost-efficient cloud inference.

Key Points

  • 2026-07-27: multiverse-computing compactifai Series C spotlight — vendor claim up to 80–95% LLM size reduction via quantum-inspired tensor networks
  • Complements quantization; academic CompactifAI work combined both on LLaMA-2
  • Central to ai-sovereignty and European efficient-AI capital narratives

Sources