Products
A short list, evaluated accurately
This is not a directory and not a marketplace. It is a curated set of infrastructure products engineering teams encounter when evaluating AI stacks, written up honestly — including where each one breaks or costs more than it looks.
Research verification status
Every product carries a badge showing the depth of our research. We do not upgrade a badge until the technical work is done — which is the only thing that makes a badge worth anything.
- Desk Research (13)
- Evaluated from public documentation, architecture material and source code. No vendor contact.
- Vendor Verified (0)
- Technical briefing completed with the vendor’s engineering team, and the claims on this page checked with them.
- Hands-On Tested (0)
- Deployed and benchmarked by us on a real workload. The notes reflect what we measured.
Today every entry is desk research. Vendor conversations are underway and the badges will move as those complete — not before. No vendor pays to be listed, ranked or recommended here.
There are 13 products here, and the test for inclusion is simple: each one is directly relevant to a problem we have written up. A product nobody has a reason to ask us about does not belong on this page yet.
Orchestration, scheduling and serving
The software deciding which job runs on which accelerator, how requests are served, and what happens when a node disappears.
vLLM
Inference serving engine
An open-source engine for serving large language models, with memory handling built for high concurrent throughput.
- Fit
- Teams self-hosting an open-weight model for production inference.
- Catch
- Serves models well; does not manage your fleet, routing or tenancy.
LiteLLM
Inference gateway
A gateway presenting one consistent interface across many hosted model providers and your own self-hosted models.
- Fit
- Applications calling more than one model provider, or planning to.
- Catch
- Adds a network hop, and sits on the critical path of every request.
SkyPilot
Multi-cloud workload orchestration
An open-source layer that runs the same job on whichever cloud, specialist GPU provider or Kubernetes cluster has capacity.
- Fit
- Teams using more than one compute provider, or wanting to.
- Catch
- You still need accounts and quota with every provider it places work on.
Run:ai
GPU pooling and fractional sharing
A Kubernetes-based orchestration layer that pools accelerators across teams and allocates fractions of one to a workload.
- Fit
- Shared clusters where several teams compete for fixed capacity.
- Catch
- Assumes you already run Kubernetes, and run it well.
Kubernetes batch scheduling
Batch scheduling
Add-ons giving Kubernetes the queueing and gang-scheduling behaviour AI workloads need and plain Kubernetes does not provide.
- Fit
- Organisations already standardised on Kubernetes.
- Catch
- You assemble a scheduler from components rather than buying one.
Compute and cloud
Where the accelerators physically are — a hyperscaler, a specialist AI cloud, or your own racks.
AWS, Google Cloud and Azure
Hyperscale cloud
The three large general-purpose clouds, each offering accelerated compute alongside everything else you already run.
- Fit
- Estates where the data, identity and compliance posture already live.
- Catch
- Generally the highest hourly rate, and quota is often the real limit.
CoreWeave
Specialist AI cloud
A cloud built specifically for accelerated workloads at scale, rather than a general-purpose cloud with GPUs added.
- Fit
- Sustained large-scale training where committed capacity makes sense.
- Catch
- Oriented to larger commitments than intermittent workloads justify.
Lambda
Specialist AI cloud
A GPU cloud aimed at AI engineering teams, which also sells hardware for on-premises deployment.
- Fit
- Teams wanting specialist pricing, and those weighing rent against buy.
- Catch
- Popular hardware types can be capacity-constrained at times.
RunPod
On-demand and serverless compute
An on-demand accelerator platform with per-second billing and a serverless mode for inference.
- Fit
- Development, experiments, and inference that scales to zero.
- Catch
- Confirm the guarantees carefully before running anything under a strict SLA.
Network and interconnect
The private links between your data centre, your clouds, your storage and everyone else.
We specify requirements and review designs at this layer; delivery is with a specialist partner.
Megaport
Network as a service
A software-defined network service for creating private connections between data centres and cloud providers on demand.
- Fit
- Hybrid estates moving enough data for transfer charges to matter.
- Catch
- Only useful where they are already present in your facility.
Equinix Fabric
Interconnection
On-demand interconnection between Equinix facilities, cloud providers and other participants in them.
- Fit
- Organisations with equipment already in an Equinix facility.
- Catch
- Compelling if you are already there; a weak reason to move.
Facility, power and cooling
The building, the power feed and the cooling a dense accelerator rack needs to run at the performance you paid for.
Evaluated from an integration and density-constraint perspective. We are not facility engineers and we coordinate specialists here.
Schneider Electric
Power distribution and facility management
Power distribution, UPS, cooling and data centre infrastructure management software.
- Fit
- Facilities being assessed or upgraded for high-density racks.
- Catch
- Facility decisions are long-lived and expensive to reverse.
LiquidStack
High-density liquid cooling
Specialist liquid cooling systems, including immersion and direct-to-chip, for rack densities air cannot handle.
- Fit
- Facilities being retrofitted or built for modern accelerator density.
- Catch
- Changes how the facility is operated, serviced and warranted.
Tracked, not yet written up
These are real, relevant and on our list. They get a page when there is a problem write-up that needs them, or a vendor conversation behind them. Publishing thin entries to look bigger would defeat the purpose of the page.
- L4 dstack
Lighter-weight alternative to SkyPilot for multi-provider orchestration.
- L4 Ray / Anyscale
Distributed compute framework for training and serving at scale.
- L4 Slurm
The established batch scheduler for owned HPC-style clusters.
- L3 Nebius
Specialist AI cloud with managed platform services.
- L3 Vast.ai
Marketplace capacity — lowest rates, highest variability.
- L2 PacketFabric
Alternative network-as-a-service provider.
- L1 Vertiv
Power and thermal management, including liquid cooling.
Something missing?
Almost certainly — this catalogue is intentionally short. If there is a product you believe belongs here, or you are a vendor interested in a technical evaluation call, tell us. There is no listing fee, no paid placement and no affiliate arrangement.
Tell us what you are trying to solve.
Describe your current setup and the problem in your own words. We reply within 2 working days with an honest assessment of whether we can help.