One and two GPU platforms for inference, fine-tuning and creative workloads. Low acoustics, standard mains power, NVMe tiers and optional 25G networking.
Quiet cooling — Tuned fan curves
1–2 GPUs — Select RTX or data centre
U.2 / U.3 NVMe — OS, scratch, datasets
GbE, optional 25G — Faster ingest
iBMC / IPMI — Remote management
About tower GPU AI servers
On-premise AI without a server room
These are for teams that need on-premise AI capability and do not have — or do not want — a dedicated server room. Meaningful GPU acceleration in a quiet chassis, with NVMe storage tiers and optional 25G networking for fast data ingest.
They run on standard office mains, with PSUs and rails sized for GPU boost draw rather than for the average. Each system is burn-in tested at sustained load, because the failure that matters in an office is the one that appears after an hour of real work.
For many teams this is the right first machine: enough capability to prove the workload, quiet enough to sit where the team sits, and a clear path to a rack node when the work outgrows it.
1 GPU — office inference and proof of concept For pilot deployments, embeddings, RAG, control interfaces and model exploration. Quiet and power-friendly. The machine that answers whether the workload is worth a rack node, without committing to one first.
2 GPU — fine-tuning and creative Twice the GPU headroom while staying quiet and office-friendly. Suited to diffusion, visual compute and small to mid-sized fine-tunes. Where a creative team needs render and preview capability locally rather than queued on a shared cluster.
Sizing examples
Three profiles, end to end
Starting points showing how the whole machine moves with the workload, not just the GPU.
Fine-tuning — GPUs: 2×. CPU: 24–32c. Memory: 128–192GB. Storage: 3× NVMe. Network: 10/25G. Notes: LoRA and QLoRA, small to mid models
Illustrative. Final sizing follows the models, frameworks and datasets you actually run.
Feature highlights
Quiet is an engineering requirement, not a preference
A GPU under sustained load produces heat that has to leave the chassis. Doing that quietly is a design problem — acoustic-tuned fan curves and ducting, rather than simply running the fans slower and losing the clocks.
The rest of the machine follows the same logic: NVMe tiers kept separate so a dataset read does not stall the OS, and images pinned so a driver update does not silently change your throughput.
Acoustic-tuned fan curves and ducting, sustaining GPU clocks while staying quiet
Dedicated NVMe for OS, scratch and datasets
Runs on standard office mains, with PSUs and rails sized for boost draw
iBMC and IPMI for remote monitoring, power cycling and updates
Driver, CUDA and cuDNN pinning for PyTorch, TensorFlow and creative stacks
profile by default
profile when you need it
standard, 10/25G optional
driver and CUDA stack
Deployment notes
What to check before it arrives
A tower in an office has constraints a rack does not.
Power — confirm the office circuit; we size rails and PSUs for GPU boost draw
Cooling — quiet profiles by default, with a performance profile available when needed
Networking — start with GbE, add 10 or 25G for faster ingest or NAS and SDS access
Storage — NVMe OS and scratch, with optional U.2/U.3 bays for dataset pools
Software — images aligned to PyTorch and TensorFlow, with hypervisors optional
Placement — quiet is designed for, but a GPU at full load is not silent; keep it off a desk surface
At a glance
The tower GPU envelope
GPUs per system
memory at the top profile
networking optional
mains power, quiet chassis
Size it first
Work the numbers before you specify
The calculators run the same arithmetic our engineers do, so you can arrive with a starting configuration rather than a blank page.
GPU & AI server sizing — Model size, batch size and dataset in; GPU count, memory, storage and fabric out.
VM density & host sizing — Host count, memory and overcommit worked through before you specify a node.
All sizing tools — Six calculators running the same arithmetic our engineers do.
Related range
Related platforms
When the work outgrows a tower.
Rack GPU AI servers — Two to six GPUs per node, with 100 or 200G fabric for distributed training.
GPU workstations — Multi-GPU compute at the desk, with high-watt PSUs and tuned thermals.
2U rack servers — The general-purpose platform underneath an AI programme.