Thermally validated, power-balanced, and prepared with driver and firmware pinning for stable AI training and inference. PCIe Gen4/Gen5, NVLink on select GPUs, U.2/U.3 NVMe and up to 200G fabric.
PCIe Gen4 / Gen5 — NVLink on select GPUs
U.2 / U.3 NVMe — OS, scratch, datasets
25 / 100 / 200G — RoCEv2 or InfiniBand
iBMC / IPMI — Telemetry and imaging
Burn-in at peak draw — GPUs under load
About rack GPU AI servers
Density you can add to, a node or a rack at a time
These nodes deliver GPU density with enterprise thermal, power and management behaviour behind it. They are built for AI workloads at scale, and sized so a programme can start small and expand by node or by rack rather than by replacement.
Every server is validated under burn-in with the GPUs at peak draw. A GPU node that passes a short benchmark and then throttles in week three is the common failure, and it is a thermal and power design problem rather than a component one.
Pick the density that matches the work: 2 GPUs for inference and fine-tuning, 4 for mainstream training, 6 where throughput per rack unit is what the budget is buying.
Three densities
Match the node to the workload
2 GPU — inference and fine-tuning — Cost-efficient nodes for RAG pipelines, embeddings and smaller fine-tunes. Suited to gateway and API layers, and to edge racks.
4 GPU — balanced training — The mainstream training node, with room for high-VRAM GPUs, NVMe pools and 100G networking. Popular LLM fine-tunes and diffusion workflows.
6 GPU — maximum density — Higher GPU count per node for accelerated training and heavy batch inference, with validated thermals, airflow and PSU headroom.
Sizing examples
Three profiles, end to end
Starting points that show how the whole node moves when GPU count changes — CPU, memory, storage and fabric all follow.
We map your framework and batch size to the right interconnect. Many deployments run excellently on PCIe alone — NVLink is a requirement, not an upgrade.
Feature highlights
What holds the clocks up in month six
A GPU node is bought on peak numbers and lived with on sustained ones. These are the parts of the design that decide whether the second figure resembles the first.
Thermals, power rails and NVMe separation do not appear on a spec sheet comparison, and they are the three things that decide whether a training run finishes at the throughput it started with.
Validated airflow, so GPUs hold boost clocks under continuous training load
High-watt PSUs sized for peak GPU draw, rail stability proven under stress
U.2 and U.3 NVMe pools, with OS and scratch kept separate
GbE through 200G RoCEv2 or InfiniBand, matched to inference or distributed training
Tool-less trays, clear cable paths and FRU documentation for fast maintenance
firmware, driver and container stack
power and thermal, out-of-band
aisle compliant airflow
high-efficiency power supplies
Deployment notes
What to plan for around the node
As much of a GPU deployment is decided outside the chassis as inside it. These are the parts to settle early.
Power — size PDUs and UPS for peak draw, not nameplate; redundant PSUs across SKUs
Cooling — hot and cold aisles with airflow paths validated for each chassis
Networking — 100 or 200G for distributed training, with RoCEv2 or InfiniBand as required
Storage — NVMe OS and scratch, with optional U.2/U.3 pools for datasets and checkpoints
Software — images aligned to PyTorch and TensorFlow, CUDA and cuDNN, drivers pinned per GPU
Expansion — plan the second node's rack position and power before the first one ships
At a glance
The rack GPU envelope
GPUs per node
memory at the top profile
fabric available
driver and firmware baseline
Size it first
Work the numbers before you specify
The calculators run the same arithmetic our engineers do, so you can arrive with a starting configuration rather than a blank page.
GPU & AI server sizing — Model size, batch size and dataset in; GPU count, memory, storage and fabric out.
VM density & host sizing — Host count, memory and overcommit worked through before you specify a node.
All sizing tools — Six calculators running the same arithmetic our engineers do.
Related range
Related platforms
Quieter, or at the desk.
Tower GPU AI servers — One or two GPUs, office-quiet, for proofs of concept and fine-tuning.
GPU workstations — Multi-GPU compute at the desk, for deep learning, path tracing and CUDA pipelines.
SNS flash storage — U.3 NVMe arrays for dataset staging and checkpoints behind the nodes.