Solution
We are about to commit, and it is hard to undo
A one to three year reservation or a hardware purchase is on the table, and the decision outlives the workload that justified it.
Reserved capacity and owned hardware are decided under genuine uncertainty. Committing buys better rates and guaranteed availability; staying flexible costs more per hour but does not tie you to a workload shape that may look completely different in two quarters. Either way, the paperwork outlives the assumptions behind it.
The failure we see most is a multi-year term sized against optimistic projections rather than measured demand, signed because the discount looked compelling in the moment. The second most common is the reverse — paying on-demand rates indefinitely for a baseline that has been stable for a year. A neutral second opinion before signature is cheap; unwinding a commitment afterwards is not.
What is usually going on
No measured baseline to size against
Committing without knowing your steady-state demand means guessing, and under a sales timeline the guess is almost always upward.
The discount is doing the deciding
A larger discount for a longer term is only a saving if you would have used the capacity anyway. Term length is a risk decision dressed up as a pricing one.
Baseline and burst treated as one thing
Most estates have a predictable floor and an unpredictable peak. Committing to the floor and bursting elsewhere is usually cheaper than either extreme.
Only one provider evaluated
Pricing, contract length and hardware availability differ substantially between hyperscalers and specialist providers. Being able to place work in more than one is a commercial position, not just a technical one.
Exit cost never examined
What it costs to leave — data egress, re-platforming, notice periods — matters as much as the headline rate, and is rarely modelled.
What we look at, in order
- 01 Establish the measured baseline and the shape of the peak, from your own telemetry
- 02 Separate what genuinely warrants a commitment from what should stay elastic
- 03 Compare providers on total cost of ownership — rate, transfer, support, and what it costs to leave
- 04 Stress-test the term against plausible changes in the workload, including it shrinking
- 05 Confirm your orchestration can actually place work with more than one provider before you rely on that as leverage
What you receive
- Demand baseline with the measurement method documented
- Commit-versus-burst allocation, with a recommended term length
- Total cost of ownership comparison across providers, including exit cost
- Risk analysis against workload change scenarios, written to be read by a budget holder
Delivered as a written assessment. More on how we work.
Products relevant to this
Not a shortlist for you specifically — that depends on constraints this page cannot know. These are the options an engineering team addressing this problem will encounter first.
CoreWeave
L3 · Specialist AI cloud
A cloud built specifically for accelerated workloads at scale, rather than a general-purpose cloud with GPUs added.
- Fit
- Sustained large-scale training where committed capacity makes sense.
- Catch
- Oriented to larger commitments than intermittent workloads justify.
Lambda
L3 · Specialist AI cloud
A GPU cloud aimed at AI engineering teams, which also sells hardware for on-premises deployment.
- Fit
- Teams wanting specialist pricing, and those weighing rent against buy.
- Catch
- Popular hardware types can be capacity-constrained at times.
RunPod
L3 · On-demand and serverless compute
An on-demand accelerator platform with per-second billing and a serverless mode for inference.
- Fit
- Development, experiments, and inference that scales to zero.
- Catch
- Confirm the guarantees carefully before running anything under a strict SLA.
AWS, Google Cloud and Azure
L3 · Hyperscale cloud
The three large general-purpose clouds, each offering accelerated compute alongside everything else you already run.
- Fit
- Estates where the data, identity and compliance posture already live.
- Catch
- Generally the highest hourly rate, and quota is often the real limit.
SkyPilot
L4 · Multi-cloud workload orchestration
An open-source layer that runs the same job on whichever cloud, specialist GPU provider or Kubernetes cluster has capacity.
- Fit
- Teams using more than one compute provider, or wanting to.
- Catch
- You still need accounts and quota with every provider it places work on.
Other problems
Tell us what you are trying to solve.
Describe your current setup and the problem in your own words. We reply within 2 working days with an honest assessment of whether we can help.