Skip to content
NuGenIT

AI compute · Orchestration · Data centre architecture

You know the problem. We help you choose the infrastructure that solves it.

NuGen IT is an independent infrastructure advisory for engineering teams running AI workloads. Whether you are costing a move from model APIs onto your own accelerators, weighing a multi-year reservation before you sign it, or trying to work out where the spend is actually going — we assess it and hand you a target architecture, a total cost comparison, and a shortlist with the trade-offs stated. We are paid by you, never by the vendors we recommend.

No cost to ask. We reply within 2 working days — and we will tell you if we are not the right people for it.

Orchestration
Deep evaluation, serving architecture and scheduling design
Compute
Deep evaluation, provider comparison and capacity strategy
Network
Network topology design and data-movement cost modelling
Facility
Density and power-constraint mapping, DCIM and specialist coordination

What are you trying to solve?

Start here rather than with a product name. Each page sets out what usually causes the problem, what we assess and in what order, and what you receive at the end.

Our inference costs rise faster than our usage

The model is in production and working. The monthly bill is growing faster than the business it supports.

L4 · L3

We are weighing model APIs against running our own

Provider invoices have grown past the point where self-hosting on dedicated accelerators is worth costing properly.

L4 · L3

It works in the pilot but not in production

The prototype proved the idea. Turning it into something dependable has stalled.

L4 · L3 · L2

We cannot see where our AI spend actually goes

The invoice arrives as one number. Nobody can attribute it to a team, a feature or a customer.

L4 · L3

We are paying for accelerators that sit idle

Reserved or owned capacity goes unused while work queues elsewhere — and preempted jobs lose their state.

L4

We are locked into one model provider

A price change, a deprecation or an outage upstream would be a serious problem for us.

L4

Our data pipeline is the bottleneck, not the compute

Expensive accelerators wait on storage and data movement instead of doing work.

L3 · L2

We need to run AI on infrastructure we control

Regulation, contractual commitments or data sensitivity mean it cannot all sit in a public cloud.

L1 · L2 · L3 · L4

We are about to commit, and it is hard to undo

A one to three year reservation or a hardware purchase is on the table, and the decision outlives the workload that justified it.

L3 · L4

Not sure which of these it is

Describe the situation in your own words and we will tell you which layer we think it sits in.

Describe it

What we actually do

You describe the problem. We come back with a written assessment — a target architecture, a cost model, a shortlist and the reasoning behind all three.

The full method
  1. 01

    You describe the problem

    Through the intake form or a short call. Tell us what is going wrong, what is slowing you down, or the decision you are stuck on — not a product name.

    No cost and no commitment. If we are not the right team for it, we say so at this point.

  2. 02

    We ask the questions that matter

    Workload type and shape, scale, latency and availability targets, what is fixed and what is negotiable, what you have already tried, and what a good outcome looks like in numbers.

    Usually one call of about 45 minutes, plus any architecture diagrams or documentation you can share.

  3. 03

    We deliver a written assessment

    A document you can circulate internally and defend in a budget conversation — not a slide deck and not a phone call you have to take notes on.

    Turnaround depends on scope and is agreed with you before we start.

  4. 04

    You decide, and we can support delivery

    Take the recommendation and run it in-house, or have us support the vendor evaluation, the proof of concept and the implementation review.

    You are never obliged to buy anything through us. We do not resell hardware, capacity or licences.

What you receive

A document you can circulate internally and defend in a budget conversation. Not a slide deck, and not a call you have to take notes on.

01

Target architecture

The recommended design across compute, orchestration, data and networking, with a diagram, and what changes from where you are today.

02

Cost model

Projected monthly cost for each option modelled side by side, with the assumptions written down so you can challenge them.

03

Shortlist and trade-off analysis

The products worth evaluating, what each one assumes you already have, and where each one breaks.

04

Risk and dependency notes

Lock-in, exit cost, single points of failure, and the operational commitment each option creates.

05

A migration sequence

What to do first, what can wait, and what should not be attempted at the same time as anything else.

06

The case for doing nothing

Where staying on your current stack is the right financial answer, that is what the document will say.

Engagement formats

Single-decision review

One specific choice, validated before you commit. A platform selection, a capacity commitment about to be signed, or a build-versus-rent question.

Short, fixed scope. Usually one call and a written recommendation.

Architecture diagnostic

Your stack assessed end to end against your workload and constraints, with a target architecture and a costed migration path.

Fixed scope and fixed fee, agreed in writing before any work starts.

Advisory retainer

Ongoing support while you deliver — vendor evaluation, design review, and a second opinion when a decision comes up.

Monthly, with an agreed scope and an agreed exit.

Scope and fee are agreed in writing before any work starts. More on pricing.

Serving and training are different problems

Much of the advice available treats "AI workloads" as one thing. They are not — and for most organisations running in production, continuous inference is now the larger and less predictable cost. The right infrastructure for one is frequently wrong for the other.

Training and fine-tuning

Interruptible, throughput-bound

  • Runs for hours or days, so being interrupted is survivable if you checkpoint properly
  • Spot and preemptible capacity is the single biggest cost lever available
  • Network between nodes decides whether you get the performance you paid for
  • Cheapest capacity anywhere usually wins, because latency to the user is irrelevant

Real-time inference

Always on, latency-bound

  • A user is waiting, so an interruption is an outage rather than an inconvenience
  • Spot capacity is usually the wrong answer here, whatever it saves
  • Where the GPU sits matters, because distance to the user is measured in milliseconds
  • Serving efficiency decides your cost — how many requests one GPU can actually handle

Where teams usually find us

Not a profile of a company — a description of a moment. If one of these is roughly where you are, the rest of the site will be useful to you.

Teams running AI in production

The system is serving real traffic. Provider invoices, accelerator utilisation, or reliability are not where they need to be, and the cost curve is steeper than the usage curve.

Teams facing a commitment they cannot easily reverse

A multi-year reservation, a hardware purchase, or a platform decision is in front of you, and you want a neutral read on the total cost before the signature rather than after it.

Teams standing up AI infrastructure for the first time

The prototype worked. Now there are questions about serving, scheduling, data access and on-call that nobody has had to answer before, and often no dedicated platform function to answer them.

Teams delivering AI systems for others

You build and ship the application. The infrastructure, networking and cost model underneath it needs depth you would rather bring in than build permanently.

Teams operating their own facilities

You are adapting existing space for accelerator density — power envelope, thermal headroom, DCIM, and the Kubernetes and GPU operator layer the workloads above will expect.

What we do not do

Worth saying plainly, because it is the part most companies leave out.

We do not resell hardware or capacity
We are not a reseller and we do not take a margin on what you buy. Our recommendation and your purchase are separate transactions.
We do not publish benchmarks we have not run
Where we cite a number, we say where it came from and when. Where we have not measured something ourselves, we say that too.
We are not facility engineers
At the power and cooling layer we map constraints and coordinate specialists. Our own evaluation depth is in serving, orchestration, compute and data movement.
We are early
We started this because the advice we wanted did not exist. We would rather tell you that than pretend to a track record we have not built.

How we stay independent

The only way this advice is worth anything

  • We are paid exclusively by our clients. Vendors do not pay us anything.
  • No vendor can pay to be listed, ranked, reviewed or recommended.
  • If we have a referral arrangement with a vendor we recommend, we say so on the recommendation.
  • We tell you when the honest answer is that you do not need us, or do not need new software at all.
More on how we work

Tell us what you are trying to solve.

Describe your current setup and the problem in your own words. We reply within 2 working days with an honest assessment of whether we can help.