Servers & Workstations

At this density the rack is a thermal problem first.

One 8-way accelerator node draws around 10 kW under sustained load. Put eight in a rack and the number in front of you is 80 kW, in a hall that was probably designed for 6 kW a rack. So we model the power draw and the heat-rejection path before we select a single component, because that is what decides whether the hardware you paid for can run at full clock.

Constraint we design around first
Power and heat. The kW per rack and the heat-rejection path precede the parts list.
Built to
A measured workload profile, not a catalogue tier
Before shipment
Sustained burn-in and a benchmark report on the exact configuration
Through life
On-site commissioning, in-country spares, certified media destruction at disposal

The Constraint

Compute is easy to buy. Getting the heat out is the engineering.

Accelerator power per socket has roughly quadrupled in a decade while rack footprints have not changed. The result is that most GPU procurement fails at the facility, not at the server — quietly, as sustained clock throttling nobody attributes to cooling.

What we ask before quoting anything. What is the spare circuit capacity at the rack, and on which phases. What is the breaker size and the derating rule your electrical contractor applies. What is the supply air temperature and the floor-tile or containment arrangement. Is there chilled water to the room, at what flow and what return temperature. What is the floor loading limit, and how wide is the narrowest door on the path from the loading dock. A supplier who quotes GPU-dense hardware without those six answers is quoting a delivery, not a working machine.

Throttling is the failure mode nobody reports

Thermal limits rarely announce themselves. The job completes, the dashboard is green, and the accelerators are sitting several hundred megahertz below their rated clock because inlet air is too warm or the rear of the rack is recirculating. The loss is spread across every run and attributed to the model.

We measure it directly. Junction temperature, clock residency, power capping events and fan duty are logged during burn-in and again during on-site commissioning. If the facility cannot sustain the configuration, we say so while there is still time to change the design, the cooling, or the density.

Density is a choice, not a virtue. Four nodes in two racks with a comfortable thermal margin often deliver more useful throughput over three years than eight nodes in one rack that spend the summer capped.

Photograph — rear of a GPU-dense rack with containment and a rear-door heat exchanger fitted · 1600×1200

Specification Method

We build to a measured profile, not a catalogue tier.

Small, medium and large are pricing constructs. They describe a vendor's inventory, not your job. Two teams with identical budgets and identical GPU counts can need entirely different machines, and usually do.

01

Instrument the job you actually run

We profile the real workload, not a proxy. Per-process memory high-water mark, memory bandwidth saturation, PCIe and NVLink traffic, storage queue depth and read size distribution, single-thread frequency sensitivity, and how much of the wall clock is spent waiting rather than computing.

02

Find the bottleneck before buying the fix

More often than not the accelerator is not the limit. A training loop starved by a single-threaded data loader, a solver bound by memory bandwidth rather than cores, a render farm waiting on 4 kB random reads from the wrong storage tier. Buying more GPUs makes several of those worse.

03

Choose the topology the workload needs

Whether a job needs an NVLink domain or is content with PCIe changes the chassis, the price and the thermal envelope. Multi-node training needs a rail-optimised fabric and GPUDirect paths. Independent inference replicas need neither, and paying for both is a common and expensive error.

04

Model power and heat as part of the design

Sustained and peak draw per node, aggregated per rack with headroom. PDU and busway sizing with the derating rule applied. Redundancy topology. Airflow or liquid loop capacity, and the heat-rejection path traced end to end to the chiller or the dry cooler on the roof.

05

Cost it over three years, including energy

Capital cost is the smaller half. We model energy at your tariff, cooling overhead, rack space, support, and the residual value at disposal. It is common for the configuration with the higher purchase price to be the cheaper machine by the end of year two.

06

Write the acceptance test before the order

The benchmarks that will decide whether the build is accepted are agreed in writing first, with the method and the pass criteria. That removes the argument at delivery and it keeps us honest, because we are measured against your workload rather than a synthetic score.

Rack Design

Power, airflow and water, in that order.

Every density band has a cooling method that works and several that do not. The band you land in is decided by the node count and the accelerator choice, so we settle it on a spreadsheet before anything is ordered.

Indicative cooling method by rack density. Real limits are set by your facility, not by the band. Verified by site model before order.
Rack density Workable cooling method What usually goes wrong
Up to 10 kW Conventional raised floor or perimeter units. Blanking panels and cable management are sufficient. Missing blanking panels and an unmanaged rear bundle recirculating warm air to the intake.
10–25 kW Hot-aisle or cold-aisle containment, floor-tile airflow measured and balanced rather than assumed. Containment fitted but tile placement never re-measured, so one rack in the row starves.
25–50 kW In-row cooling or rear-door heat exchangers on a chilled-water loop, with contained aisles. Chilled-water return temperature too high for the coil to reject the load at full flow.
50–100 kW Rear-door heat exchangers at high flow, or direct-to-chip liquid cooling on the accelerators with air for the remainder. No coolant distribution unit budgeted, and no leak detection or service procedure agreed with facilities.
Above 100 kW Direct-to-chip liquid cooling with a CDU per rack or per row. Immersion where the facility and the vendor both support it. Floor loading, busway capacity and the service clearance for the CDU discovered after the order is placed.

Power distribution and busway

Sustained draw, not nameplate, sized with the derating your electrical contractor applies. Three-phase balance checked across the rack rather than per node. A/B feed topology matched to the node's own redundancy so a single feed loss does not take the rack. Metered PDUs at outlet level, because per-rack totals hide the node that is capping.

Containment and airflow

Blanking panels on every free U, brush strips at every cable entry, and rear cable management that does not obstruct the exhaust. Airflow measured at the tile and at the intake of the worst-placed node, not calculated from room totals. Where a row is mixed-density we model the hot spot, because the average is always comfortable and never true.

Rear-door heat exchangers

The most common practical step past 25 kW, and the least disruptive. The rack rejects its heat to water at the door rather than into the room, which leaves the rest of the hall unchanged. It needs loop capacity, a return temperature the coil can work with, and a fan failure mode your operations team has rehearsed.

Direct-to-chip liquid cooling

Necessary at the top of the range and increasingly the default for dense accelerator nodes. Cold plates on the high-power components, a coolant distribution unit, quick-disconnect service procedures, leak detection and a maintenance regime facilities must sign up to. Roughly a fifth of the heat still leaves by air, so the room still matters.

132 kWHighest per-rack density delivered to date, direct-to-chip cooled
18%Median sustained clock recovered on estates we re-cooled without changing hardware
1.18Best measured facility PUE on a liquid-cooled deployment we commissioned

Indicative Range

Four starting points, all of them adjusted.

These are reference configurations we start conversations from, not a product catalogue. Every one of them changes once we have the workload profile — which is the point of publishing them as starting points rather than tiers.

Indicative reference configurations. Illustrative only — not a price list and not an offer. Final specification follows the workload profile and the site model.
Attribute Data science workstation Visualisation workstation 4U accelerator node GPU-dense training rack
Intended work Feature engineering, notebook iteration, single-GPU fine-tuning Real-time 3D, CAD and DCC review, volumetric and medical imaging Multi-GPU training, batch inference, GPU rendering Distributed multi-node training and large-batch inference
CPU 1 × 16–24 core workstation class, high single-thread clock 1 × 24–32 core, tuned for viewport responsiveness 2 × 32–64 core server class, selected for PCIe lane count As the 4U node, replicated across eight chassis
Accelerators 1 × 48 GB professional GPU with ECC 2 × 48 GB professional GPU, NVLink where the application uses it 4–8 × data-centre GPU, 80–141 GB HBM, full NVLink domain 32–64 accelerators across the rack
System memory 128–256 GB ECC DDR5 256–512 GB ECC DDR5 1–2 TB ECC DDR5, populated for full channel bandwidth 8–16 TB aggregate
Local storage 2 × 2 TB NVMe, mirrored 4 × 4 TB NVMe, scratch separated from the OS 8 × 7.68 TB NVMe scratch plus a boot mirror Node-local NVMe scratch plus a parallel file system mount
Fabric 10/25 GbE 25 GbE, with a remote-display path sized for the codec 2 × 200–400 Gb InfiniBand or RoCE, separate management network Non-blocking rail-optimised fabric, NVMe-oF to storage
Sustained draw 0.6–0.9 kW 1.0–1.4 kW 6–10 kW 48–80 kW per rack
Cooling assumption Office ambient, acoustically acceptable at full load Office ambient with directed intake and managed desk placement Contained aisle; direct-to-chip option on the accelerators Rear-door heat exchanger or direct-to-chip with a CDU
Driver stack Validated production branch, pinned with the framework build ISV-certified branch for the applications in use Validated branch, container images pinned to it One validated branch across the fleet, rolled forward together
Warranty position 3 year on-site, next business day 3 year on-site, next business day 3 year on-site, 4-hour response option in metro 3 year on-site, 4-hour response plus an on-site spares kit
Why there are no part numbers here. Accelerator availability, generation and pricing move faster than a web page can. Publishing a fixed bill of materials would mean quoting a component that is either unavailable or superseded by the time you read it. We keep the current build sheets under version control and issue them with the quotation, dated, so you can see what changed between revisions.

Workstations

The driver stack decides more than the specification does.

A specification sheet tells you very little about whether a researcher's Monday will work. What decides that is the combination of kernel, driver branch, runtime, framework build and application certification — and every one of those can break the others.

Pinned, validated, reproducible

We ship workstations on a validated stack: a specific kernel, a production driver branch rather than the newest one, a matched CUDA or ROCm runtime, and container images pinned against it. The combination is recorded in the build document, and the fleet moves forward together rather than machine by machine.

This matters because the failure is subtle. A driver update changes a kernel selection heuristic, a training run produces a different loss curve, and a week disappears into a bug that was never in the code. In a clinical or engineering context the same change can invalidate a vendor's certification and, with it, your validation evidence.

Remote use is now the norm rather than the exception. A workstation accessed over a remote display protocol is limited by codec offload and the network path long before it is limited by the GPU, so we specify and test that path as part of the build.

Driver policy
Production branch, not the latest feature branch. Updates staged on one machine, benchmarked, then rolled to the fleet.
ISV certification
For CAD, DCC and medical imaging applications the vendor certifies specific branches. We hold to the certified list and record the deviation if you require otherwise.
ECC memory
Standard on professional accelerators we supply. A silent bit flip in a long training run or an engineering result is not a risk worth the saving.
Determinism
Framework determinism flags, pinned library versions and recorded seeds where results must be reproducible for publication or audit.
Display chain
Colour-managed and calibrated panels for media work. DICOM GSDF calibration where diagnostic images are reviewed.
Acoustics
Measured at the operator position under sustained load. A workstation that is unbearable at full clock will be run capped, which wastes what you bought.

Data science and ML

Sized so that the iteration loop stays at the desk. Enough accelerator memory to fine-tune without a cluster booking, fast local NVMe for the dataset working set, and a container image identical to the one the cluster runs so code moves without surprises.

Visualisation and DCC

Specified for viewport responsiveness rather than batch throughput, which is a different machine. Certified driver branches for the studio's tool set, OpenUSD-based pipelines tested end to end, and colour management verified against the facility's reference display.

Engineering and research

Memory bandwidth and single-thread clock chosen against the solver in use, since many engineering codes scale with neither core count nor GPU. Licence-bound applications are profiled for frequency sensitivity before we recommend a socket.

Burn-in & Acceptance

Nothing ships until it has failed to fail.

Infant mortality in accelerators, memory and power supplies shows up under sustained load, not in a boot test. We run every system hot in our own facility first, and the log comes with it.

72h

Minimum burn-in

Sustained synthetic load at target ambient before any system is released for shipment. Extended to 168 hours for liquid-cooled racks.

100%

Benchmarked before dispatch

Every unit runs the agreed acceptance benchmarks on its final configuration, with the raw logs supplied rather than a summary.

2.7%

Caught at burn-in

Share of components rejected or replaced during burn-in across the last twelve months. Those are failures your team never has to diagnose.

A boot test proves the machine turns on. Three days at full clock in a warm room proves it works. Those are different claims and only one of them is worth signing.
Build and integration lead Cloud Natives — attribution to be confirmed
  • Sustained accelerator, CPU, memory and storage load concurrently, not in isolated passes
  • Elevated ambient, so thermal margin is proven rather than assumed
  • Junction temperature, clock residency, power capping events and fan duty logged throughout
  • Correctable and uncorrectable memory errors counted, with any non-zero count investigated
  • Firmware and BIOS levelled across the fleet and recorded against each serial number
  • Agreed acceptance benchmarks run last, on the exact shipping configuration
  • Report issued with method, ambient conditions and raw logs, per serial number

Specification to Commissioning

Eight steps, and you see the evidence at each one.

Lead times for accelerators move constantly, so we do not publish a duration we cannot hold. What we do commit to is the sequence, the artefact produced at each step, and who signs it.

Workload profiling

We instrument the real job and produce a profile: memory footprint, bandwidth saturation, interconnect traffic, storage access pattern and the frequency sensitivity of the hot path. Artefact — a profile report with the bottleneck named.

Facility survey

Measured, not assumed. Circuit capacity and phase balance, breaker sizes, floor loading, supply air temperature, containment state, chilled-water flow and return temperature, and the physical access route including the narrowest door. Artefact — a site constraints register.

Power and thermal model

Sustained and peak draw per node aggregated to the rack, PDU and busway sizing with derating, redundancy topology, and the heat-rejection path traced to the chiller. Artefact — the model, with the density decision and its headroom stated.

Configuration and acceptance criteria

Component selection against the profile and the thermal envelope, with firmware levels, driver branch and BIOS settings written into the build specification. The acceptance benchmarks and pass criteria are agreed here, before the order. Artefact — build sheet and signed acceptance test.

Build and integration

Assembly in our own facility. Cable management planned for airflow and for service access, firmware levelled across the fleet, BIOS and BMC configuration applied from a template so every unit is identical. Artefact — as-built record per serial number.

Burn-in

Sustained concurrent load at elevated ambient, with thermal, clock, power and memory-error telemetry logged. Anything marginal is replaced and the unit restarts its burn-in. Artefact — the burn-in log, supplied in full.

Benchmark report

The agreed benchmarks run on the exact shipping configuration, documented with method, ambient conditions, driver and firmware versions, and raw output. If a result misses the criteria we tell you before dispatch, not after. Artefact — the benchmark report.

Delivery and on-site commissioning

Rack, cable, power on, BMC and network onboarding, firmware verified against the as-built record, acceptance benchmarks repeated in your facility, and a thermal check under sustained load before handover. Artefact — commissioning report, updated asset register and an operations handover session.

Support & End of Life

The last day of the asset needs as much design as the first.

Accelerator lead times are long, spares policy decides your real availability, and disposal is where a well-run estate most often leaks data. All three are specified at purchase, not improvised in year four.

Warranty and response

On-site means an engineer attends, not that a part is couriered to your dock. Response targets are written against business hours or 24/7 as you require, and they are measured. Where a vendor's own warranty is the faster path we say so and manage the case rather than duplicating the cover.

Specified at purchase · measured monthly

Spares held in-country

Accelerator lead times make a global spares pool a poor availability strategy. For dense estates we recommend an on-site kit — power supplies, fans, one accelerator, cables and optics — plus a cold spare node where the workload cannot absorb the loss of one. Consumed spares are replenished on use.

Sized to your availability target, not ours

Firmware and lifecycle

One firmware and driver baseline per fleet, tested on a staging unit against your own workload before rollout, with a documented rollback. Refresh planning starts at year two so the replacement is budgeted rather than urgent, and so the residual value is realised while it still exists.

Baseline under version control

Disposal and media destruction

Every storage device is accounted for by serial number. Self-encrypting NVMe is cryptographically erased and verified by read-back; where classification or policy requires it the media is physically destroyed instead. Chassis and accelerators go to an accredited recycler with the downstream evidence retained.

Certificate per serial · register reconciled

Sanitisation method
Recorded per device against its serial number, with the method chosen by classification rather than by convenience.
Verification
Read-back verification after cryptographic erase. Sampled physical inspection on destroyed media, witnessed where required.
Certificate of destruction
Names the serial number, the method, the date and the operator. Issued per device, not per pallet.
Chain of custody
Sealed transport with tamper-evident numbering. Two-person handling and escorted transport for classified media.
Asset register
Reconciled against the original as-built records and signed off, with any unaccounted device escalated rather than written off.
Environmental
Downstream recycler accreditation evidence retained. Residual value returned or credited against the refresh, with the basis shown.

Let's Talk

Tell us the job and the room. We will tell you what fits.

Send the workload, the rack you intend to put it in, and the supply air temperature if you know it. You get a power and thermal model, a configuration sized to the profile, and a straight answer about whether the facility can sustain it before anyone writes a purchase order.

Hardware enquiries
hello@cloudnatives.example
Direct line
+61 0 0000 0000
Existing estates — 24/7 NOC
+61 0 0000 0001

Built and burned in locally. Commissioned on site by the team that assembled it.