Definition
Inference chips are specialized semiconductors designed to run AI model inference (predictions) efficiently, as opposed to training chips used for model development. Q1 2026 saw a major shift toward inference chip investments.
Landmark: Google Frozen v2 (Reported, July 2026)
- 2026-07: Reported frozen-v2 hardwires gemini architecture with updatable weights; 6–10× tokens/watt claim; ~2028 (google-frozen-v2-gemini-chip)
- Parallel to TPUs; model-specific path under model-specific-inference-silicon
Warning
Unverified The Information leak — Google has not confirmed.
The Inference Shift
The semiconductor industry is experiencing an “inference shift”:
- Previous focus: Training GPUs (H100, H200) for model development
- Current focus: Specialized inference chips for deployment
- Drivers: Cost efficiency, latency, power consumption at scale
Landmark: Etched Sohu Stealth Exit (June 2026)
- 2026-06-30: etched emerged from stealth with working sohu-chip on TSMC N4P — 1B+ customer contracts (2026-06-30-etched-800m-inference-chip-stealth)
- Transformer attention hardcoded into silicon; inference-only ASIC bet
- First racks shipping summer 2026; gigawatt-scale production target 2027
Warning
Etched throughput/cost claims are company-reported; independent benchmarks pending.
Landmark: OpenAI Jalapeño (June 2026)
- 2026-06-24: openai and broadcom unveiled jalapeno-chip — first custom OpenAI inference ASIC
- 9-month design-to-tape-out; AI-assisted chip design; ~50% cost savings vs GPUs (early testing, Broadcom CEO)
- Inference-only; deployment in microsoft data centers targeted end 2026
- See jalapeno-chip for full specifications
Landmark: SambaNova SN50 (July 2026)
- 2026-07-08: sambanova 11B valuation — General Atlantic lead (2026-07-08-sambanova-1b-series-f-11b-valuation)
- jpmorgan on-premises inference partner; softbank first SN50 customer
- SN50 ships H2 2026; five months after $350M Series E
- Signals inference hardware as next capital battleground after training chips
Q1 2026 Funding Landscape
| Company | Funding | Focus |
|---|---|---|
| cerebras | $1B Series H | AI training/inference |
| rapidus | $1.7B | 2nm chips by 2027 |
| matx | $500M | LLM inference |
| etched | $500M | LLM inference |
| kandou | $225M | Interconnect technology |
Key Trends
- Specialization: Chips purpose-built for transformer architectures
- Optical interconnects: For data center efficiency
- Power management: Specialized semiconductors for AI workloads
- Geographic diversification: India emerging as VC hub
Industry Metrics
- AI chips generate 50% of semiconductor revenue despite only 0.2% of unit volume
- Global semiconductor revenue projected: $975 billion in 2026
- nvidia still dominant but faces new competition