Definition
Model-specific inference silicon embeds a particular neural-network architecture (and sometimes weights) into hardware, trading flexibility for tokens-per-watt and latency versus general-purpose GPUs/TPUs.
Key Points
- 2026-07: Reported frozen-v2 hardwires gemini architecture with updatable weights (google-frozen-v2-gemini-chip)
- 2026-06: etched sohu-chip hardcodes transformer attention for inference-only ASICs
- 2026-06: openai jalapeno-chip with broadcom — custom inference ASIC path
- Core trade-off: architecture lock-in vs efficiency under compute-scarcity
- Fits custom-ai-silicon trend away from pure GPU rental