Specifying compute for an edge site with no engineer on it

An unattended edge computer must be remotely diagnosable and locally replaceable. The specification should cover cabinet inlet temperature, dust, condensation, remote management, power recovery and site-held spares, not only processor performance.

An edge site without a resident engineer should be specified as a remotely diagnosable, locally replaceable system, not as a maintenance-free computer. Processor performance matters, but continued operation depends equally on the cabinet, cooling path, storage, power supply, UPS, network and management connection. The practical objective is to detect a developing fault remotely, place the system in a safe state where necessary, identify the replaceable item and give a local technician an unambiguous replacement procedure. Specify the installed operating envelope The operating temperature printed on a processor or compute-module data sheet is not the operating envelope of the installed machine. Every item in the cabinet has its own limits, including storage devices, power supplies, carrier boards, network switches, optical modules, UPS batteries and display or camera interfaces. The system must be designed around the narrowest applicable limit. This is also the approach recommended in ASHRAE guidance for edge facilities . Measure temperature where the machine receives its air A tender statement such as ambient temperature up to 50°C is incomplete unless the measurement point is defined. The cabinet air presented to the compute-system inlet may be considerably hotter because of solar heating, restricted circulation, batteries, power conversion equipment and adjacent electronics. The site survey and technical specification should record: Minimum and maximum temperature at the compute-system air inlet or specified chassis reference point. Temperature with the processor, GPU and storage under maximum sustained load. Temperature reached after loss of room air-conditioning or cabinet cooling. Rate of temperature change during start-up, shutdown and cabinet opening. Relative humidity, dew point and likelihood of condensation. Installation altitude. Direct and reflected solar exposure. Heat released by the UPS, batteries, network equipment and other cabinet loads. Acceptance testing should use sustained workload rather than a short boot or diagnostic test. CPU temperature, GPU temperature, storage temperature, fan speed, throttling status and cabinet inlet temperature should be logged together. Calculate cooling for the complete cabinet Passive cooling is usable only when sufficient heat can pass from the cabinet interior to the surrounding environment. If the outside air is already at or above the permitted internal temperature, a sealed passive enclosure cannot reject the compute load while maintaining that internal limit. The thermal calculation should include continuous equipment dissipation, solar gain, enclosure surface area, allowable temperature rise, filter loading, fan ageing and cooling-capacity reduction at high ambient temperature. Natural convection, filtered ventilation, air-to-air heat exchange and active cooling are different engineering arrangements, not interchangeable descriptions. Rittal's enclosure cooling guidance gives a useful overview of these methods. An outdoor cabinet may require a sun shield, double-skin roof or shaded mounting position. These are parts of the thermal system. The bidder should state whether the quoted maximum operating temperature assumes shade or includes the declared solar load. Treat condensation as a separate condition A humidity range marked non-condensing does not demonstrate suitability for every field condition. Condensation can occur when a warm, humid cabinet cools after a power failure, when cold equipment is exposed to humid air, or when a door is opened during servicing. Possible controls include a hygrostat, anti-condensation heater, controlled ventilation and a restart delay until temperature and humidity stabilise. IEC 60068-2-78 addresses steady-state damp heat without condensation. Temperature-change testing is covered separately by IEC 60068-2-14. Cite the edition current on your bid date rather than a remembered one. The tender should identify which condition is relevant rather than citing a general humidity percentage alone. State the altitude Air cooling becomes less effective as air density falls. Component limits and any temperature derating must therefore be checked at the actual installation altitude. A sea-level thermal result should not be accepted automatically for a hill site. Dust protection must include the completed assembly IEC 60529 defines the IP code used for enclosure protection. The required rating should apply to the completed arrangement with doors, cable glands, vents, filters, antenna entries and external connectors installed. An IP-rated empty cabinet does not establish the rating after field modifications. Sealing also restricts airflow. A specification cannot independently demand a highly sealed cabinet, unrestricted air cooling and high internal heat dissipation without defining how the heat crosses the enclosure boundary. Choose a controlled cooling path For a dusty site, the design should state which of these arrangements is being supplied: A sealed, conduction-cooled or fanless computer transferring heat to the enclosure or mounting surface. A sealed cabinet using a closed-loop air-to-air or air-to-water heat exchanger. A positively pressurised cabinet with replaceable intake filters and monitored airflow. An actively cooled sealed cabinet sized for the declared ambient temperature and heat load. Fanless does not mean thermally unrestricted. A fanless embedded computer still requires the specified mounting orientation, clearance and heat-transfer path. Thermal performance should be verified in the final cabinet, not only on an open laboratory bench. Make filters maintainable A filter is a consumable item. As it loads with dust, airflow falls and internal temperature rises. The filter grade, dimensions, part number, replacement interval and access method should be documented. If replacement requires removal of live power wiring or complete cabinet dismantling, routine maintenance is unlikely to happen at the intended interval. Useful monitoring includes cabinet differential pressure, fan tachometer status and temperature before and after the filter. Where these sensors are not provided, a conservative inspection interval should be established from site experience. Specify compute form factor against the site Unattended edge installations commonly use a fanless embedded computer, a compact GPU edge appliance, or a short-depth 1U or 2U server. The correct choice depends on cabinet depth, sustained power, environmental control and the permitted service method. Fanless embedded systems A fanless system removes one moving component and can reduce dust circulation through the compute chassis. It is suitable only when the selected CPU, memory, storage and accelerators remain within their limits at the declared mounting surface and ambient temperature. Short-depth rack systems A short-depth 1U or 2U system can provide server-class remote management, ECC memory and replaceable storage or power supplies. For many edge workloads, a single-socket platform is easier to cool than a dual-socket server. Dual-socket compute should be specified only where the workload, memory capacity or I/O requirement justifies the additional power and heat. Memory should be stated as an installed configuration, not only as a maximum platform capacity. For example, specify 64 GB or 128 GB ECC memory, the number and capacity of fitted DIMMs, available slots, and whether memory sparing or mirroring is required. Jetson Orin edge AI systems A Jetson Orin installation must be assessed as the module, carrier board, storage, power input, enclosure and software image together. The compute module's temperature range alone does not qualify the completed appliance. The specification should identify the exact Jetson Orin module, configured power mode, sustained AI workload, memory capacity, storage type, camera interfaces and carrier-board operating limits. Remote recovery facilities vary by carrier board and appliance design, so they must be verified for the quoted system rather than assumed from the module name. Suitable platform options can be reviewed under Jetson Orin edge AI systems . Remote management must work when the operating system does not Remote desktop software is not out-of-band management. It is unavailable when the operating system has stopped, the boot device has failed or the network stack is misconfigured. For server-class systems, specify a dedicated BMC with an independent management interface. Required functions should include: Remote power on, controlled shutdown, reset and power cycle. Remote console access from power-on self-test onwards. BIOS or UEFI configuration and remote virtual media. Hardware event logs and sensor readings for temperature, voltage, fans and power supplies. Role-based access, encrypted protocols and auditable login records. Redfish or another documented interface for integration with the monitoring platform. Configuration backup and a controlled firmware update procedure. Compact embedded and Jetson systems may not contain a server BMC. In that case, equivalent recovery functions must be designed explicitly. These can include an independent hardware watchdog, a remotely controlled power distribution unit, recovery partitions, automatic rollback after a failed software update and a serial console through a separate management device. Keep the management path independent A remotely switchable outlet is of little use if it depends on the failed edge computer for connectivity. The management path should survive failure of the workload operating system. Depending on site criticality, this may require a separate management VLAN, an independent router or cellular connection, and separately powered cabinet monitoring. Remote management access should not be exposed directly to the public Internet. Access control, certificate management, VPN policy, log retention and credential rotation should be part of deployment documentation. Define power-failure behaviour The tender should state what happens after mains failure, UPS exhaustion and restoration of supply. Required settings may include automatic power-on after AC recovery, delayed restart to avoid simultaneous inrush, graceful shutdown on UPS low-battery status and watchdog recovery after a software hang. Test the complete sequence with the actual UPS and power controller. A BIOS setting alone does not prove recovery if the UPS remains latched off or the network device starts later than the compute system. Plan for the components most exposed to wear and environment There is no universal first failure. In unattended cabinets, faults often appear first in items affected by heat, dust, vibration, repeated switching or limited write endurance rather than in the processor itself. The maintenance plan should pay particular attention to: Air filters blocked by dust. Fans with worn bearings or reduced speed. UPS batteries exposed to elevated temperature. Boot storage subjected to frequent logging or uncontrolled shutdowns. Power supplies affected by heat, unstable input power or dust. Cable glands, connectors and network links exposed to moisture, vibration or corrosion. Cabinet cooling units and condensate arrangements. Storage writes should be controlled. Operating-system logs, application telemetry, video buffers and AI inference records can create a continuous write workload. Specify enterprise SSDs or suitable industrial storage with declared endurance, expose SMART or NVMe health data to monitoring, and reserve sufficient free capacity. A mirrored pair improves availability against a device failure but does not replace backup or protect against software corruption. Use a site-held spares strategy A four-hour diagnosis is not useful if the replacement part requires several days to reach a remote location. Spares should be selected by failure impact, replacement difficulty, installed population and lead time. Typical site or regional spares include: Pre-imaged boot SSD or complete storage set. Matched power supply or external power adapter. Fan and filter kit. UPS battery cartridge where local replacement is permitted. Configured network switch, router or communication modem. Complete compute node for sealed or tightly integrated appliances. Approved cables, glands and connector assemblies. For a modular 1U or 2U server, component-level replacement may be practical. For a sealed fanless unit or compact Jetson appliance, swapping the complete pre-configured node can reduce site time and avoid opening the enclosure in dust or humidity. Every spare should have a controlled software image, firmware baseline, configuration record and periodic test procedure. An untested spare stored for years is not an assured recovery method. Write acceptance criteria that can be demonstrated An edge-compute tender should convert general requirements into measurable conditions. Useful clauses include: Maximum declared temperature at the compute inlet under sustained CPU, GPU, storage and network load. Operation at the stated altitude, humidity condition and cabinet heat load. Evidence for the IP rating of the final enclosure arrangement. Maximum permitted CPU or GPU throttling during the acceptance workload. Remote console, remote reset and virtual-media demonstration with the operating system unavailable. Automatic recovery test following mains failure and UPS shutdown. Alarm generation for over-temperature, fan failure, storage degradation and loss of external power. Replacement demonstration for filters, storage and the nominated field-replaceable unit. Submission of wiring diagrams, port lists, firmware versions, software image checksums and recovery instructions. The final specification should be based on the actual site survey and workload. NetBytes industrial and edge systems can be configured around fanless embedded, rack-mounted and accelerated edge requirements. For workload sizing, cabinet integration and remote-management planning, share the site conditions through our contact page .

Topics: edge computing, industrial computers, remote management, thermal design, dust protection, edge AI, Jetson Orin, unattended sites

More articles · Talk to our engineers