This page may contain stale information. Last updated: 2026-04-21
Definition
NPU (Neural Processing Unit) is specialized hardware designed for efficient inference of neural networks. Distinct from general-purpose GPUs, NPUs are optimized for low-power, on-device AI operations like language models and computer vision.
Key Characteristics
- Dedicated to neural networks: Fixed-function hardware for matrix operations
- Low power consumption: Enables mobile/edge deployment (watts vs. kilowatts)
- On-device processing: Privacy-preserving local computation
- Inference-focused: Optimized for inference, not training
Performance Metrics
Measured in TOPS (Tera Operations Per Second). Examples:
- amd-ryzen-ai-400: up to 60 TOPS (mobile), 50 TOPS (desktop)
- qualcomm-snapdragon-x: ~45 TOPS
- Apple Neural Engine: ~16 TOPS (but specialized for Apple ML models)
copilot-plus-pc Requirements
Microsoft’s copilot-plus-pc standard requires:
- Minimum 40 TOPS NPU
- Enables on-device AI features (local Copilot, image generation)
- amd, intel (Meteor Lake), qualcomm competing for market
Market Emergence (2025-2026)
NPU integration into consumer chips represents major shift:
- Previously: AI required cloud servers or discrete GPUs
- Now: 40-60 TOPS on every new PC/phone
- Implications: Local model inference, privacy, reduced cloud dependency