Compare 2U, 3U and 4U GPU servers by qualified accelerator capacity, drive layout, cooling geometry, chassis depth and actual rack power limits.
A 2U GPU server commonly supports up to four double-width PCIe accelerators, while a purpose-built 4U server may support eight and, in specialised designs, ten. A 3U server is not automatically a six-GPU system: the additional height may instead accommodate an SXM baseboard, larger fans, more drives, PCIe switches or liquid-cooling hardware. The correct choice is therefore not the chassis with the largest stated slot count. It is the smallest qualified configuration that accepts the required accelerators, sustains their power and cooling requirements, fits the available rack, and leaves usable capacity for storage and network adapters. This comparison covers conventional 2U, 3U and 4U systems. It does not cover the entire GPU-server market. High-power eight-GPU HGX platforms may occupy 6U, 8U or more, especially where large air-cooled baseboards and high-capacity power systems are required. What 2U, 3U and 4U actually specify One rack unit is 44.45 mm high. The nominal vertical envelopes are therefore 88.90 mm for 2U, 133.35 mm for 3U and 177.80 mm for 4U. Actual chassis dimensions are slightly smaller to provide installation clearance. The rack-unit designation specifies height, not depth, GPU count, drive capacity or electrical load. EIA-310 defines the 19-inch mounting arrangement and vertical rack-unit system, but buyers must separately verify chassis depth, rail adjustment, cable clearance and rack load capacity. A useful overview of rack dimensions is available from Vertiv . What a 2U GPU server buys High accelerator density per rack unit A purpose-built 2U server can accommodate four full-height, full-length, double-width PCIe accelerators. Representative dual-socket systems combine four PCIe 5.0 x16 GPU positions with additional slots for network adapters, DPUs or storage controllers. Some also support paired GPU interconnect bridges. Four physical GPU positions do not prove that every accelerator can be installed four at a time. A chassis may qualify four lower-power cards but only two 600 W cards because of power, connector, airflow or component-temperature limits. The tender should specify the exact GPU model and required quantity as one qualified population. Buyers considering this density class can review 2U rack server configurations . Storage can be compact, but it is not always limited GPU-focused 2U systems commonly provide around eight front NVMe bays in E1.S, E3.S or 2.5-inch formats. This is generally adequate for mirrored boot media, local cache, model staging and scratch data. However, 2U does not impose a universal eight-drive limit. Some general-purpose 2U architectures combine four double-width GPU positions with as many as 22 front 2.5-inch bays. Whether such a combination is usable depends on the motherboard orientation, backplane, risers, fan wall and the qualified GPU power class. Drive count should therefore be evaluated together with media type, PCIe lane allocation and inlet obstruction. Eight direct-attached NVMe bays and 22 SAS or SATA bays represent different storage and thermal designs even if both systems are 2U. Cooling is concentrated Four passive accelerators in 2U place a substantial heat load behind a front inlet only two rack units high. The design depends on high-pressure fans, air shrouds, sealed airflow paths, correct blanking panels and the thermal kit qualified for the installed cards. A representative four-GPU 2U platform is 900 mm deep and uses four 2,000 W Titanium-rated power supplies. PSU nameplate wattage indicates the capacity and redundancy class; it is not the measured consumption of every configuration. The platform details are published by Supermicro . Where 2U is the practical choice One to four qualified PCIe accelerators are required per node. GPU count per rack unit is more important than extensive internal storage. The rack and cooling system can handle concentrated heat at each 2U position. Cluster growth is expected through more nodes rather than eight or ten GPUs in one node. The required NICs, DPUs and storage adapters remain available after the GPUs are fitted. What a 3U GPU server buys 3U is an architecture choice, not a midpoint A 3U GPU server may still provide four accelerators. Its extra rack unit can be used for a dedicated accelerator baseboard, PCIe switches, larger fans, drive bays, network expansion or direct liquid cooling rather than two more double-width cards. One four-GPU 3U architecture combined H200 SXM accelerators and NVLink with two AMD EPYC sockets, 24 DIMM slots, eight front 2.5-inch bays and six low-profile PCIe 5.0 x16 positions. The direct-liquid-cooled variant used three 3,000 W power supplies in a 2+1 arrangement, measured approximately 850.2 mm deep and weighed about 50.4 kg. The cited model is now marked as legacy, so it is an architectural example rather than a current purchase recommendation. Its specifications remain available from GIGABYTE . Current 3U availability and the exact accelerator qualification should be confirmed for the configuration under offer. Buyers can review the applicable 3U rack server range and obtain a configuration statement before fixing tender parameters. The additional height can support different cooling methods Chassis height does not reveal whether a server is air cooled or liquid cooled. Closely related 3U designs have been offered with both air-cooled and direct-liquid-cooled four-GPU configurations. Direct liquid cooling also does not necessarily remove all server fans. CPUs, DIMMs, NICs, drives, voltage regulators and power supplies may still require forced air. Coolant temperature, flow, water quality, condensation control and facility connection requirements become part of the server specification. GPU, DPU and NIC qualifications must be read together. A high-power network card can alter the permitted ambient temperature or fan requirement even when the accelerators themselves are liquid cooled. 3U can favour drives and conventional expansion A general-purpose 3U chassis can provide seven full-height, full-length expansion positions, eight 3.5-inch hot-swap bays and two 5.25-inch bays at a depth of about 647 mm. This does not make it a qualified multi-GPU server, but it shows what the height can provide when storage and standard PCIe expansion take priority over maximum accelerator density. Where 3U is the practical choice A four-GPU SXM or specialised baseboard architecture is required. The design needs more cooling cross-section than a dense 2U node. Additional drive, fabric or low-profile expansion positions are valuable. A qualified direct-liquid-cooled configuration matches the facility cooling system. The selected 3U platform has a clear lifecycle and support status for the tender period. What a 4U GPU server buys Eight double-width GPUs become practical A purpose-built 4U PCIe GPU server can support eight full-height, full-length, double-width accelerators. Some specialised dual-root and PCIe-switch designs support up to ten, subject to the qualified accelerator list and system power limits. Ten mechanical positions should not be converted into a tender requirement without checking CPU-to-GPU topology, peer-to-peer communication, PCIe switch arrangement and supported GPU wattage. An eight-GPU configuration with the required topology may be preferable to a ten-slot chassis that does not support the intended workload. Relevant platforms can be reviewed under 4U rack servers . Drive-bay capacity depends on the internal layout There is no standard drive count for a 4U GPU server. Current and recent designs include eight rear E1.S NVMe bays, eight front NVMe bays with sixteen SAS or SATA bays, and arrangements offering up to 24 front 2.5-inch bays. A specialised ten-GPU platform may provide only two SATA and eight NVMe positions. Bay placement matters as much as quantity. A large front drive backplane consumes inlet area and adds airflow resistance. Rear E1.S bays preserve more of the front inlet for GPU cooling, but they affect rear service access, cable routing and hot-aisle working space. More height improves cooling geometry A 4U chassis provides a larger fan cross-section, wider airflow channels and more scope to separate GPU, CPU, storage and PSU air paths. Representative eight-GPU systems use fan walls containing up to ten 80 mm fans. The extra height does not remove the need for GPU qualification. Passive server accelerators depend on chassis airflow. For reference, the AMD Instinct MI210 is a full-height, full-length, double-slot passive PCIe card rated at 300 W, while the NVIDIA L40S is a full-height, full-length, dual-slot PCIe card rated at up to 350 W. Their form factor and cooling method must match the server rather than merely its available slot count. Power becomes a rack-level constraint One current eight-GPU 4U design uses four 3,200 W Titanium-rated power supplies in a 3+1 arrangement, with full rated PSU output requiring 220 to 240 V input. Other 4U designs use four 2,000 W supplies. These figures describe PSU capacity and redundancy, not a guaranteed operating draw. The bid should state maximum input power for the offered CPU, GPU, memory, drive and NIC population. Typical workload consumption can be supplied separately, but it should not replace the maximum figure used for electrical planning. Where 4U is the practical choice Eight qualified PCIe accelerators are required in one node. GPU-to-GPU locality within a server is more important than maximum node count per rack. The configuration needs a larger fan wall and less restrictive airflow paths. More internal drives or fabric adapters are required alongside the GPUs. The rack can support the chassis depth, installed weight, input power and rear cabling. GPU form factor must be stated separately PCIe slot count is not GPU capacity “Eight PCIe slots” does not mean that eight double-width GPUs can be installed. The specification should state whether each accelerator is single-slot, double-slot or wider, together with full-height, low-profile, full-length or half-length dimensions. It should also state the electrical lane width and generation. A mechanically x16 slot may be electrically connected as x8, and adjacent connectors may be blocked by a double-width card. Auxiliary power connector type and available power per position must also be confirmed. PCI-SIG lists PCI Express Card Electromechanical Specification Revision 6.0.1, dated 13 March 2025, as the approved CEM specification. A tender purchasing PCIe 5.0 GPUs does not normally need to mandate that CEM revision by itself. It should specify the required PCIe generation, lane width, card envelope and power support. The specification status is available from PCI-SIG . SXM and OAM are not ordinary add-in cards SXM and OAM accelerators are installed on dedicated baseboards. Their node height is determined by the complete baseboard, interconnect, cooling and power architecture rather than by conventional PCIe slot spacing. A four-SXM node can fit in 3U, while high-power eight-GPU HGX systems commonly use larger chassis. For example, an eight-GPU H100 or H200 platform can occupy 6U at nearly 995 mm depth, while the NVIDIA DGX H100/H200 occupies 8U, is approximately 897.1 mm deep and has a specified maximum power of 10.2 kW. NVIDIA publishes the DGX dimensions and power information in its system documentation . Rack depth can eliminate a configuration Height is a poor predictor of chassis depth. Representative examples include a 900 mm-deep 2U GPU server, an 850.2 mm-deep 3U accelerator server, a 737 mm-deep 4U GPU server and a 6U HGX system approaching 995 mm. The rack check must include more than nominal chassis depth. Verify the following: Front-post to rear-post distance and the rail adjustment range. Clearance behind the chassis for power connectors, network cables and fibre bend radius. Space occupied by vertical PDUs and their sockets. Rear-door clearance and the effect of high cable density on exhaust airflow. Front bezel, cold-aisle containment and door perforation. Installed weight, rail rating and safe service access. A 900 mm server should not be assumed to fit properly in a nominal 1,000 mm rack. Connector bodies, cable bends, PDUs and door clearance can consume the remaining space. Power per rack usually sets the real density A 42U rack has space for twenty-one 2U servers, fourteen 3U servers or ten 4U servers before reserving space for switches and other equipment. A GPU rack will often reach its electrical or cooling limit well before these physical counts. For example, if an approved rack envelope is 30 kW and the declared planning load is 6 kW per server, only five such servers fit within the power envelope even if ten or more fit by height. This is an arithmetic example, not a specification for a particular NetBytes system. Rack planning should account for: Maximum configured input power, not only the average AI workload. A and B feed capacity under the required redundancy policy. Whether the rack must remain operational after one PSU or one feed fails. PDU socket type, branch-circuit rating and phase balance. 220 to 240 V input requirements for full PSU output. Cooling capacity in kW per rack and the permitted inlet-temperature range. Any site derating required by the electrical design or operating practice. High GPU density per U has little value if alternate rack positions must be left empty for power or cooling reasons. In such a site, 4U nodes may provide similar usable GPUs per rack with simpler cabling and less concentrated airflow than densely packed 2U nodes. Storage frontage affects accelerator cooling Front drive bays compete with fans and GPUs for inlet air. Dense 2.5-inch, E3.S or 3.5-inch backplanes create resistance before the air reaches the accelerator zone. This is particularly important for passive cards, which do not carry their own cooling fans. Drive requirements should distinguish among boot, local cache, scratch space and retained datasets. A GPU node may need only mirrored boot drives and a small NVMe scratch tier when training data resides on external storage. Another workload may require substantial local NVMe capacity to avoid network bottlenecks. Do not specify the maximum possible bay count merely as a higher numerical requirement. Ask for the installed media, usable capacity, RAID or software-defined layout, endurance class and measured path to the GPUs or CPUs. Useful tender requirements Define the offered accelerator population GPU make and exact model. Quantity installed and maximum qualified quantity. PCIe, SXM or OAM form factor. Card width, height, length and thermal design power. Required GPU interconnect and topology. Confirmation that all GPUs operate concurrently without unsupported power capping. Define the complete server configuration Chassis height and exact depth without and with bezel. Processor socket count and offered processors. DIMM slot count, installed memory and maximum supported memory. Drive form factors, interface, installed quantity and remaining bays. Available PCIe slots after all GPUs, NICs, DPUs and storage controllers are fitted. Rail kit range, cable-management arrangement and maximum configured weight. Define power and cooling evidence PSU quantity, rating, efficiency grade and redundancy mode. Required input voltage and connector type. Maximum input power for the offered configuration. Supported ambient inlet-temperature range. Air-cooled or direct-liquid-cooled configuration. For liquid cooling, coolant, flow, pressure, temperature, connection and condensation-control requirements. Manufacturer qualification for the complete GPU, NIC, memory and drive population. A practical selection rule Select 2U for one to four PCIe GPUs where rack power and airflow are already controlled; select 3U for a qualified architecture that uses the additional height for baseboard, cooling, storage or fabric; select 4U when eight PCIe GPUs, wider airflow paths or greater internal expansion are required. The comparison should be made using bill-of-material-level configurations. NetBytes GPU and AI systems are designed, built, burn-in tested and supported in India, and selected systems also operate in our own data centre. Available configurations can be reviewed under products , with deployment guidance in the knowledge base . For rack, power and accelerator qualification against a tender requirement, share the complete specification through contact .
Topics: GPU servers, 2U rack servers, 3U rack servers, 4U rack servers, AI infrastructure, data centre planning, server procurement