AI compute · Orchestration · Data centre architecture
You know the problem. We help you choose the infrastructure that solves it.
NuGen IT is an independent infrastructure advisory for engineering teams running AI workloads. Whether you are costing a move from model APIs onto your own accelerators, weighing a multi-year reservation before you sign it, or trying to work out where the spend is actually going — we assess it and hand you a target architecture, a total cost comparison, and a shortlist with the trade-offs stated. We are paid by you, never by the vendors we recommend.
No cost to ask. We reply within 2 working days — and we will tell you if we are not the right people for it.
- Orchestration
- Deep evaluation, serving architecture and scheduling design
- Compute
- Deep evaluation, provider comparison and capacity strategy
- Network
- Network topology design and data-movement cost modelling
- Facility
- Density and power-constraint mapping, DCIM and specialist coordination
What are you trying to solve?
Start here rather than with a product name. Each page sets out what usually causes the problem, what we assess and in what order, and what you receive at the end.
Our inference costs rise faster than our usage
The model is in production and working. The monthly bill is growing faster than the business it supports.
L4 · L3We are weighing model APIs against running our own
Provider invoices have grown past the point where self-hosting on dedicated accelerators is worth costing properly.
L4 · L3It works in the pilot but not in production
The prototype proved the idea. Turning it into something dependable has stalled.
L4 · L3 · L2We cannot see where our AI spend actually goes
The invoice arrives as one number. Nobody can attribute it to a team, a feature or a customer.
L4 · L3We are paying for accelerators that sit idle
Reserved or owned capacity goes unused while work queues elsewhere — and preempted jobs lose their state.
L4We are locked into one model provider
A price change, a deprecation or an outage upstream would be a serious problem for us.
L4Our data pipeline is the bottleneck, not the compute
Expensive accelerators wait on storage and data movement instead of doing work.
L3 · L2We need to run AI on infrastructure we control
Regulation, contractual commitments or data sensitivity mean it cannot all sit in a public cloud.
L1 · L2 · L3 · L4We are about to commit, and it is hard to undo
A one to three year reservation or a hardware purchase is on the table, and the decision outlives the workload that justified it.
L3 · L4Not sure which of these it is
Describe the situation in your own words and we will tell you which layer we think it sits in.
Describe itWhat we actually do
You describe the problem. We come back with a written assessment — a target architecture, a cost model, a shortlist and the reasoning behind all three.
The full method- 01
You describe the problem
Through the intake form or a short call. Tell us what is going wrong, what is slowing you down, or the decision you are stuck on — not a product name.
No cost and no commitment. If we are not the right team for it, we say so at this point.
- 02
We ask the questions that matter
Workload type and shape, scale, latency and availability targets, what is fixed and what is negotiable, what you have already tried, and what a good outcome looks like in numbers.
Usually one call of about 45 minutes, plus any architecture diagrams or documentation you can share.
- 03
We deliver a written assessment
A document you can circulate internally and defend in a budget conversation — not a slide deck and not a phone call you have to take notes on.
Turnaround depends on scope and is agreed with you before we start.
- 04
You decide, and we can support delivery
Take the recommendation and run it in-house, or have us support the vendor evaluation, the proof of concept and the implementation review.
You are never obliged to buy anything through us. We do not resell hardware, capacity or licences.
What you receive
A document you can circulate internally and defend in a budget conversation. Not a slide deck, and not a call you have to take notes on.
Target architecture
The recommended design across compute, orchestration, data and networking, with a diagram, and what changes from where you are today.
Cost model
Projected monthly cost for each option modelled side by side, with the assumptions written down so you can challenge them.
Shortlist and trade-off analysis
The products worth evaluating, what each one assumes you already have, and where each one breaks.
Risk and dependency notes
Lock-in, exit cost, single points of failure, and the operational commitment each option creates.
A migration sequence
What to do first, what can wait, and what should not be attempted at the same time as anything else.
The case for doing nothing
Where staying on your current stack is the right financial answer, that is what the document will say.
Engagement formats
Single-decision review
One specific choice, validated before you commit. A platform selection, a capacity commitment about to be signed, or a build-versus-rent question.
Short, fixed scope. Usually one call and a written recommendation.
Architecture diagnostic
Your stack assessed end to end against your workload and constraints, with a target architecture and a costed migration path.
Fixed scope and fixed fee, agreed in writing before any work starts.
Advisory retainer
Ongoing support while you deliver — vendor evaluation, design review, and a second opinion when a decision comes up.
Monthly, with an agreed scope and an agreed exit.
Scope and fee are agreed in writing before any work starts. More on pricing.
Serving and training are different problems
Much of the advice available treats "AI workloads" as one thing. They are not — and for most organisations running in production, continuous inference is now the larger and less predictable cost. The right infrastructure for one is frequently wrong for the other.
Training and fine-tuning
Interruptible, throughput-bound
- Runs for hours or days, so being interrupted is survivable if you checkpoint properly
- Spot and preemptible capacity is the single biggest cost lever available
- Network between nodes decides whether you get the performance you paid for
- Cheapest capacity anywhere usually wins, because latency to the user is irrelevant
Real-time inference
Always on, latency-bound
- A user is waiting, so an interruption is an outage rather than an inconvenience
- Spot capacity is usually the wrong answer here, whatever it saves
- Where the GPU sits matters, because distance to the user is measured in milliseconds
- Serving efficiency decides your cost — how many requests one GPU can actually handle
One team across all four layers
Power and thermal limits at the bottom, the scheduler and serving engine at the top, and the interconnect and capacity decisions in between. Most problems turn out to live in one specific layer — naming which one is half the work, and it is hard to do that if you can only see part of the stack.
Orchestration, scheduling and serving
Deep evaluationThe software deciding which job runs on which accelerator, how requests are served, and what happens when a node disappears.
Compute and cloud
Deep evaluationWhere the accelerators physically are — a hyperscaler, a specialist AI cloud, or your own racks.
Network and interconnect
Architecture and designThe private links between your data centre, your clouds, your storage and everyone else.
Facility, power and cooling
Requirements and constraintsThe building, the power feed and the cooling a dense accelerator rack needs to run at the performance you paid for.
Where teams usually find us
Not a profile of a company — a description of a moment. If one of these is roughly where you are, the rest of the site will be useful to you.
Teams running AI in production
The system is serving real traffic. Provider invoices, accelerator utilisation, or reliability are not where they need to be, and the cost curve is steeper than the usage curve.
Teams facing a commitment they cannot easily reverse
A multi-year reservation, a hardware purchase, or a platform decision is in front of you, and you want a neutral read on the total cost before the signature rather than after it.
Teams standing up AI infrastructure for the first time
The prototype worked. Now there are questions about serving, scheduling, data access and on-call that nobody has had to answer before, and often no dedicated platform function to answer them.
Teams delivering AI systems for others
You build and ship the application. The infrastructure, networking and cost model underneath it needs depth you would rather bring in than build permanently.
Teams operating their own facilities
You are adapting existing space for accelerator density — power envelope, thermal headroom, DCIM, and the Kubernetes and GPU operator layer the workloads above will expect.
What we do not do
Worth saying plainly, because it is the part most companies leave out.
- We do not resell hardware or capacity
- We are not a reseller and we do not take a margin on what you buy. Our recommendation and your purchase are separate transactions.
- We do not publish benchmarks we have not run
- Where we cite a number, we say where it came from and when. Where we have not measured something ourselves, we say that too.
- We are not facility engineers
- At the power and cooling layer we map constraints and coordinate specialists. Our own evaluation depth is in serving, orchestration, compute and data movement.
- We are early
- We started this because the advice we wanted did not exist. We would rather tell you that than pretend to a track record we have not built.
How we stay independent
The only way this advice is worth anything
- We are paid exclusively by our clients. Vendors do not pay us anything.
- No vendor can pay to be listed, ranked, reviewed or recommended.
- If we have a referral arrangement with a vendor we recommend, we say so on the recommendation.
- We tell you when the honest answer is that you do not need us, or do not need new software at all.
Tell us what you are trying to solve.
Describe your current setup and the problem in your own words. We reply within 2 working days with an honest assessment of whether we can help.