Up to 2.8x GPU, 3.7x energy, 4.3x cost reduction while maintaining SLOs. Profile-guided optimizer + adaptive runtime.