News - cracker method for AI cost optimisation?
AI costs are splitting in two directions. Microsoft is cutting cloud prices while GPU and inference bills explode. This article shows why “AI infrastructure” is no longer one line in your budget, and how that blind spot …
Connections
AI-assisted with Frank's editing.
Yesterday's hyperscaler news that caught my eye is a Microsoft one, and it made me think about how cloud pricing is fracturing in ways that break our forecasting assumptions.
Microsoft just announced 7.4% price cuts for European commercial cloud services, effective February 2026 [1]. GPU prices, meanwhile, are set to rise 20-30% this year due to HBM memory shortages [2].
Strategic moves, both of them. And they don't contradict each other at all.
I like that Microsoft is explicit about this. They publish the rationale: currency alignment with USD, applied transparently across EUR, CHF, and SEK [1]. AWS, on the other hand, prices everything in dollars and leaves currency risk to customers. One approach feels helpful; the other protects margin.
Cloud pricing has fractured. That's the real story here.
Commodity services (CPU, basic networking, storage) face competitive pressure, so vendors cut prices. Specialised AI infrastructure (GPUs, high-bandwidth memory) faces supply constraints and limited alternatives, so vendors raise prices. According to Astute Group, HBM now represents over 80% of the bill of materials for high-end GPUs, and memory prices have risen several hundred percent in recent months. SK Hynix and Micron have already sold out their 2026 HBM production capacity [2].
But here's the deeper split that matters for forecasting: training versus inference.
Training costs are becoming predictable and finite. ByteIota reports that GPT-4's training cost was roughly $150M [3]. Large, yes. But bounded. This is where in-house GPUs start making strategic sense—own the hardware, capitalise the investment, depreciate over time. Training might take slightly longer on your own infrastructure, but you gain cost certainty and control.
Inference is the opposite. It's continuous, variable, and scales with demand. That same GPT-4 model accumulated $2.3B in inference costs by end of 2024 [3]. For every $1B spent training an AI model, organizations face an estimated $15-20B in inference costs over the model's production lifetime [3]. This is where cloud makes sense: pay for what you use, scale elastically, no stranded capacity when models evolve.
The pattern emerging: train in-house when costs stabilise, move inference to cloud when GPU prices spike, then potentially migrate back when the market rebalances. It's a divergent-convergent strategy. If you own GPUs, planning for easy migration between on-prem and cloud becomes strategic infrastructure work. I expect we'll see new technologies built exactly for this.
I'd suggest forecasting components separately. Training GPU-hours, inference GPU-hours, CPU, storage, transfer. Apply different price trajectories to each based on market structure.
And if your forecast still treats "AI infrastructure" as one line, you're probably underestimating inference by 5-10x.
Sources: