Definition
Techniques that reduce model size, memory, and compute cost — including quantization, pruning, distillation, and tensor-networks — to enable edge-ai, on-prem, and cost-efficient cloud inference.
Key Points
- 2026-07-27: multiverse-computing compactifai Series C spotlight — vendor claim up to 80–95% LLM size reduction via quantum-inspired tensor networks
- Complements quantization; academic CompactifAI work combined both on LLaMA-2
- Central to ai-sovereignty and European efficient-AI capital narratives
Related
- compactifai
- multiverse-computing
- tensor-networks
- quantization
- edge-ai
- european-efficient-ai-capital
- ai-inference