A short functional check proves that a machine works at one point in time. Burn-in keeps the completed system under sustained activity so that marginal components, assembly defects and integration problems have an opportunity to appear before dispatch.
Every NetBytes machine is burned in before dispatch because sustained operation can reveal early-life and integration defects that a point-in-time functional check may not expose. The machine is tested after final assembly, in the configuration ordered by the customer. A burn-in period is not a guarantee against every future fault. It is a production screen intended to reduce the risk of early field failure and provide a failure-free operating record before dispatch. This forms part of how NetBytes systems are designed, built, tested and supported in India, as described on Why NetBytes . Functional testing and burn-in answer different questions A functional test confirms that a machine works at the time of inspection; burn-in checks whether the fully assembled machine continues to work correctly under sustained load, temperature and repeated activity. A functional test remains necessary. It confirms that the system powers on, completes POST, recognises the ordered processors and memory, detects storage and network devices, and boots the required operating environment. It can also verify ports, management access and basic application operation. However, a machine that passes these checks has not necessarily reached stable temperatures or experienced repeated changes in electrical load. A marginal DIMM contact, an intermittently operating fan or an unstable PCIe link may remain quiet during a short test. Burn-in gives these faults more time and operating conditions in which to become visible. Why infant mortality deserves factory time Infant mortality is the initially elevated failure rate associated with latent component defects, workmanship errors, assembly variation and manufacturing-process deviations. The term does not mean that every component is expected to fail early. It describes a population-level pattern in which weak units tend to fail nearer the beginning of service. The failure rate then declines as those early defects are removed. The NIST reliability handbook describes this declining early-life region of the commonly used bathtub curve. For an enterprise machine, an early failure is disproportionately disruptive. It can occur during OS deployment, cluster commissioning, data migration, application validation or site acceptance. The resulting cost is not limited to replacing a part. It may also involve engineer travel, repeat testing, tender documentation, downtime and delayed handover. Burn-in moves part of that risk into the factory, where the complete system can be inspected and corrected before packing. It cannot cover an infant-mortality period that may extend for weeks or months, but it can precipitate some defects that depend on operating time, heat or repeated transitions. What a 24-to-72-hour burn-in can catch There is no universal standard requiring every enterprise machine to undergo exactly 24, 48 or 72 hours of burn-in. The suitable duration depends on the product class, component count, power density, ordered configuration and acceptance requirement. Duration alone is also not sufficient. A machine left powered on but idle has undergone little meaningful screening. A useful burn-in combines sustained workload, changing workload, health monitoring and inspection of hardware event records. The following are the principal fault classes such a run can expose. Cooling and thermal assembly faults A processor heat sink can be slightly mis-seated while still allowing the machine to boot. Thermal-interface material may be uneven, a fan may operate intermittently, or internal cabling may restrict airflow. These conditions become clearer after processors, GPUs, memory and drives have generated heat for a sustained period. The relevant evidence is more than the absence of an over-temperature shutdown. Test records should show available inlet, processor and GPU temperatures or thermal margins, fan speeds, power consumption, throttling events and maximum readings. BMC and Redfish records can provide standardised health, temperature, fan, power-supply and event information on supported platforms. Marginal DIMMs and memory population A memory fault may require a particular data pattern, address range or operating temperature before it appears. Longer pattern testing can expose marginal DIMMs, weak slot contacts and population problems that a boot-time memory count does not detect. This matters on one-socket and two-socket servers with memory distributed across several channels and NUMA nodes. A configuration with 16, 24 or 32 DIMMs should exercise memory attached to every populated processor, not merely allocate a small block on the first node. ECC also requires careful interpretation. A corrected error may not stop the operating system, but it is still evidence requiring investigation when it first appears during production testing. Linux EDAC, BMC logs and platform diagnostics can distinguish corrected events from uncorrected errors. PCIe, GPU and accelerator-path instability GPUs, high-speed NICs, RAID controllers, HBAs and NVMe devices depend on the complete PCIe path through processors, risers, connectors and system firmware. Hardware can recover from some PCIe errors without an immediate application failure. The event may be visible only as an AER record, replay count, link retraining event or reduced negotiated link state. GPU burn-in must therefore do more than run a sample programme. Appropriate testing exercises framebuffer memory, sustained compute, host-to-device transfers and peer paths where applicable. Event inspection should cover GPU ECC errors, Xid records, PCIe replays, NVLink health, thermal excursions and unstable performance. NVIDIA DCGM diagnostics , where supported, include active tests for GPU memory, compute, PCIe, power and thermal behaviour. This is distinct from passive monitoring that reports only whether a recent health alert has occurred. Power-delivery faults under changing load A constant full load checks sustained power and cooling capacity. Rapid changes between lower and higher load can expose a different class of problem involving power supplies, voltage regulation or accelerator boards. This is why a useful test profile contains both steady and changing workloads. Where redundant PSUs are ordered, detecting both supplies is not the same as proving redundancy. A controlled check should confirm PSU health and the required behaviour after removal of one supply or one input feed, subject to the approved procedure. The tender or QAP should state the expected result clearly. Drive, backplane and storage-controller defects A storage volume can be created successfully even when one drive, connector or backplane path is marginal. Burn-in should account for every installed drive rather than checking only the logical volume presented by the RAID controller. Depending on the configuration, checks can include NVMe self-tests, SMART health, media errors, drive temperature, controller cache status, RAID initialisation, hot-spare recognition and path events. HDD long tests may take several hours because they examine far more of the media surface than a quick functional check. A controlled rebuild or failover test can be relevant for storage arrays, but it must be defined against the ordered RAID level and controller arrangement. SSD workloads must also have a sensible write budget. Consuming flash endurance merely to fill elapsed time is not sound testing. Intermittent assembly and connection problems Processors, DIMMs, drives and add-in cards may have passed their respective manufacturers' tests. The completed machine can still contain an integration fault involving a riser, cable, backplane, connector, retention mechanism or firmware combination. Heating, cooling and load changes cause small mechanical and electrical variations across these assemblies. Extended operation provides more opportunities for an intermittent connection to generate a detectable error. This is why testing the actual ordered machine and serial number is materially different from testing one reference configuration. Restart and firmware-state faults A continuous load does not exercise every platform transition. Warm reboots and cold starts test processor and memory initialisation, storage-controller discovery, device enumeration and boot behaviour from a different direction. Relevant checks include BIOS and BMC event records, management access, recognition of every configured CPU, DIMM, drive, NIC and accelerator, and PXE or network boot where it forms part of the ordered requirement. Repetition can expose a device that appears on some starts but not others. Burn-in must be instrumented, not merely timed The value of burn-in comes from component coverage, workload and recorded evidence, not from elapsed hours alone. A practical production record should identify the tested serial number and its actual hardware configuration. Depending on the machine, useful evidence includes: Processor count, sustained utilisation and maximum temperature or thermal margin. Installed memory capacity, DIMM count, channel population and new ECC events. GPU count, compute workload, framebuffer test result, temperature and Xid or ECC records. PCIe negotiated link state, AER events, replay or retraining information where available. Every installed HDD or SSD, its health status, temperature and relevant media errors. RAID, HBA and backplane status, including cache and path health. PSU health, fan speeds and power readings available from platform management. Warm-restart and cold-start results, including device recognition after restart. BIOS, BMC, operating-system and controller logs reviewed at the end of the run. A new corrected ECC event or recoverable PCIe error should not be dismissed merely because the workload continued. Recoverable faults are often the evidence that distinguishes a marginal assembly from a clean one. Burn-in is not environmental qualification Routine production burn-in must stay within the operating limits of the platform and its components. Excessive processor temperature, unrestricted SSD writes or stress beyond datasheet limits can consume useful service life rather than improve screening. Burn-in also does not prove compliance with vibration, shock, humidity, thermal-shock or ingress requirements. IEC 60068 treats these as separate environmental tests. Compliance with IEC 60068, IS 9000, JSS 55555 or a specific defence requirement must be supported by the applicable type-test or qualification evidence. Similarly, BIS registration for automatic data-processing equipment addresses the applicable safety standard. It should not be read as certification of a particular reliability burn-in period, MTBF or field-failure rate. How to specify burn-in in a tender or QAP A requirement stating only “72-hour burn-in” leaves important questions unanswered. For meaningful and comparable compliance, the tender or QAP should define: Whether the duration is total elapsed time or a failure-free operating interval. The workload applied to CPU, memory, GPU, network and storage subsystems. Whether testing is at normal factory ambient or a specified controlled temperature. The monitoring points and logs to be retained. Permitted corrected errors, if any, and conditions requiring investigation or replacement. Warm-reboot, cold-start, PSU failover and storage-rebuild requirements where applicable. Whether repair, firmware change or hardware reconfiguration restarts the required interval. The serial-number-wise report or certificate required with dispatch documents. The workload should reflect the ordered machine. A dual-socket server with 32 DIMMs requires different coverage from a tower workstation. A multi-GPU AI system requires accelerator memory, compute, PCIe and power-transition testing that does not apply to a basic file server. Burn-in reduces risk; warranty remains necessary Burn-in cannot find every random future failure, nor can it reproduce every rack, power, workload and environmental condition at the deployment site. It does not replace correct site preparation, redundancy, preventive maintenance, spares or support. Its purpose is narrower and practical: identify some marginal components and assemblies before the machine enters service. Faults that occur later remain subject to the applicable support terms. NetBytes warranty information is available at Warranty , while product and configuration enquiries can be submitted through Contact .