Managed Services

Managed services, and the engineer who answers at 3 a.m.

Operations is where infrastructure value is either realised or quietly lost. We monitor, patch, tune and cost-control the estates we build — from a network operations centre staffed in Australia, by engineers who were on the design review.

The Constraint We Design Around First

Who answers at 3 a.m.

Every managed services proposal promises 24/7. The question that decides the outcome is who picks up. And whether they can already see your architecture, or are reading it for the first time while the cluster is down.

A build ends. Operations does not. The value of a well-specified cluster is realised across the four years after commissioning, or lost in them one week at a time.

Most degradation is not an outage. It is a queue that drains a little slower each week. A storage tier whose p99 has doubled since March. A GPU clocked down by a failing fan, still billing at full rate. None of that trips a reachability check, and none of it appears in a monthly uptime figure.

So we staff operations with the people who did the design. Tier three escalation reaches an engineer who holds the as-built topology in their head, not a triage script written by someone who has never seen your fabric. That is the whole differentiator. It is also why we will not operate an estate we were not permitted to document.

The responder already knows the architecture, because we built it.
The first operating rule for every estate we take on. Everything else on this page exists to make it true at 3 a.m.

Operations Centre

A NOC with the design team standing behind it.

In-country staffing, four escalation tiers, and a monitoring scope that goes well past reachability. We manage alert quality, never alert volume.

Operations floor photography — 1600×1000

Sovereign NOC. Shift roster, console wall and escalation board.

Every shift is worked from Australia. There is no follow-the-sun handover to an offshore desk at 11 p.m., because a handover is where context dies and where an incident stops making progress.

  • Australian-staffed shifts, with a named personnel list maintained under your contract
  • Console access via jump hosts inside your jurisdiction, session-recorded end to end
  • Break-glass credentials held in your vault, checked out against a ticket, rotated on use
  • Cleared personnel only on PROTECTED estates, with access events shipped to your SIEM

Monitoring telemetry stays in the same jurisdiction as the estate it describes. We do not ship metrics to a shared offshore observability tenancy, because for government and defence work that single decision undoes the rest of the design.

Escalation tiers

Tier 1 — Operations analyst

On shift, in-country. Owns acknowledgement, classification, runbook execution and the incident record. Instructed to escalate rather than improvise.

Holds the incident end to end

Tier 2 — Platform engineer

Scheduler, fabric, storage and hypervisor depth. Holds root on your estate and the change authority to use it inside an agreed window.

Diagnosis and remediation

Tier 3 — Design engineer

The engineer who specified the build. Reached by name from your escalation card, not through a queue. Joins any P1 by default rather than on request.

Architectural authority

Tier 4 — Vendor and facility

Silicon vendor support contracts, colocation remote hands and the data centre's own operations desk. We drive those cases and keep the incident open on our side.

We hold the thread, not you

Telemetry

Up or down is the least interesting signal an estate produces.

These are the series we alarm on, and the reason each one matters weeks before it becomes an outage. Thresholds are published to you and any of them can be vetoed.

Scheduler and queue health

Slurm queue depth by partition, pending reason codes, fair-share drift between projects, and job failure rate by node. A queue that lengthens for three straight days is a capacity decision, not an incident — and it should reach a human as exactly that.

Leading indicator — capacity exhaustion

GPU health and throttling

Corrected and uncorrected ECC counts, retired pages, NVLink and PCIe replay errors, and SM clock measured against the thermal and power caps. A throttled accelerator costs the same per hour as a healthy one and finishes nothing on time.

Leading indicator — silent performance loss

Fabric integrity

InfiniBand symbol-error and link-downed counters, congestion and credit-stall metrics, RDMA retransmit rates, and port flap history per leaf. One marginal cable degrades every collective operation that crosses it, and the job log blames the application.

Leading indicator — one bad optic or cable

Storage latency distribution

Per-target latency percentiles across Lustre and BeeGFS, NVMe-oF namespace queue depth, metadata operation rates, and rebuild or scrub progress. We alarm on p99 and p99.9. An average read latency has never once told anyone the truth.

Leading indicator — tail latency before collapse

Thermal and power headroom

Rack inlet temperature and delta-T, coolant distribution unit supply temperature and flow rate, PDU phase balance and branch-circuit headroom, fan and pump duty cycles. Headroom is the number that predicts a February outage in November.

Leading indicator — summer capacity limits

Time and network discipline

PTP offset and path-delay variance, PPS signal integrity, grandmaster holdover state, plus BGP session health and egress volume per flow. Trading and timestamping estates fail on time discipline long before they fail on compute.

Leading indicator — timestamp and audit risk

Alert quality over alert volume

A NOC that pages on everything trains its own staff to ignore it. Within a month the operator who acknowledges fastest is the one who reads least. That failure mode is cultural, it is predictable, and it is created by the monitoring configuration rather than by the people.

So every alarm we create carries three things: a runbook, a named owner, and a review date. If an alarm fires three times without a decision attached to it, the alarm is wrong. We rewrite it or delete it, and the change appears in your monthly service report with the reasoning.

3 Median actionable alerts reaching a human per estate, per night
96% Alarms resolved by documented runbook without escalation
0 Standing alarms accepted as known noise at handover

Engagement Models

Three tiers. Choose against your own bench strength.

Monitor if you have capable engineers and want a second set of eyes. Operate if you want the work done. Fully Managed if you want one accountable party for the platform, its documentation and its spend.

Indicative scope only. Response targets are design planning figures and are superseded in full by the signed services agreement.
Service dimension Monitor Operate Fully Managed
Coverage hours 24/7 monitoring, response in business hours (AEST/AEDT) 24/7/365 monitoring and response 24/7/365, with tier three on call
P1 response target 30 minute acknowledgement, advisory only 15 minute acknowledgement, engineer engaged within 30 minutes 15 minute acknowledgement, engineer engaged within 15 minutes
P2 response target Next business day 4 business hours 2 hours, around the clock
Remediation authority You act, we advise and document We act under runbooks you have approved We act, including unscripted diagnosis and recovery
Patching and firmware Advisory notices, you schedule the work Quarterly OS updates, monthly security patching in your window Monthly OS, quarterly firmware and BIOS baseline, canary node first
Capacity review cadence Quarterly written report Monthly written review of utilisation and queue trend Monthly review plus a quarterly capacity and commitment plan
On-site attendance By the hour, on request Next business day for hardware faults 4 hour metro target, next available flight regional, local spares held
Named engineer Shared operations queue Named platform engineer Named platform engineer and named design engineer
Runbook ownership You write them, we review annually We write and maintain them We write, maintain and game-day test them quarterly
Change execution Yours entirely Ours, in your window, with minuted change approval Ours, against a change freeze calendar we hold on your behalf
Reporting Monthly availability and alert summary Monthly service report with incident register Monthly report plus a quarterly service review with your executive
Read this before quoting any number above. Response targets shown on this page are indicative design figures used for planning conversations. Real service levels — the measurement point, the exclusions, the escalation obligations and the service credits — exist only in a signed services agreement. If a target matters to your business case, ask for it in writing and we will either commit to it or explain why we will not.

Deployment & Handover

We will not hand over an undocumented system.

Commissioning is not the moment it powers on. It is a written acceptance test against the numbers in the proposal, run in front of your team, reporting the misses as plainly as the passes.

Sample acceptance test report — 1200×1600

Measured result beside proposed result, line by line.

Acceptance testing re-runs the benchmark suite named in the proposal, on your data where licence and privacy allow. It prints the measured number beside the proposed one. Where we fall short we state by how much and why, before you sign.

  • Burn-in and thermal soak under synthetic load before the estate carries real work
  • Functional acceptance — scheduler, authentication, quotas, failover, backup restore
  • Performance acceptance — the quoted benchmark suite, re-run on the installed hardware
  • Security baseline — hardening checklist, patch level, credential rotation, log shipping

A restore is tested during commissioning, not described. Until a file has come back from the backup tier and been checksummed, you do not have a backup — you have an expense with optimistic documentation.

The handover pack

As-built
Rack elevations, cabling schedule, IP and VLAN plan, firmware baseline, and a bill of materials with serial numbers.
Runbooks
Start, stop, drain, patch, fail over and restore. Each one executed at least once with your team watching before we call it done.
Acceptance record
Test plan, measured results, and the variances you accepted — with the name of the person who accepted each one.
Knowledge transfer
Two structured sessions and a recorded walkthrough, delivered before the final payment milestone rather than after it.
Escalation card
Names, numbers and tiers on one page, printed for the wall and stored where an on-call engineer can find it without a login.
Spares register
What is held, where it is held, and the replacement lead time for each line item including optics and power supplies.
When documentation competes with a date, we move the date. An undocumented estate becomes unsupportable within about a year — usually the week after the person who built it changes jobs. We have declined handover sign-off over this, and we would do it again. It is cheaper for everyone than the alternative, and it is the only way the 3 a.m. promise on this page survives contact with staff turnover.

Cloud Optimisation

Most overspend is not the rate. It is idle non-production capacity.

We have yet to review an estate where the largest recoverable line was a badly negotiated price. It is nearly always compute nobody switched off — development, staging, and a training cluster from a project that finished last year.

01

Rightsize against observed utilisation, not the original ticket

We take 90 days of CPU, memory, GPU and IOPS telemetry per workload and size to the measured p95, with headroom you agree to in writing. Instance families shift generation to generation, so we re-test rather than assume the newest silicon is cheaper for your particular shape. Memory-bound workloads routinely get worse on a faster core count.

02

Commitment strategy — reserved capacity versus savings plans

Reserved instances win where the shape is fixed and the term is survivable. Savings plans win where the shape moves but the spend does not. We model both against your own twelve-month usage curve, buy in tranches rather than one decision, and hold a deliberate on-demand margin. A commitment you have grown out of becomes a floor you are paying to ignore.

03

Egress and data-path control

Egress is the cost that never appears in the sizing spreadsheet. We map every cross-region, cross-account and internet-bound flow, then put the chatty ones behind a cache or a private interconnect. After that we move the workload to the data, rather than the data to the workload. For estates with a sovereignty obligation, the same map doubles as evidence of where the bytes actually go.

04

Storage class and lifecycle policy, written per dataset

Class policy comes from access telemetry, not habit. Hot tier for the working set, infrequent access past thirty days, archive on a retention rule, deletion on a legal one. We test the retrieval before we trust the tier. An archive you cannot restore inside your recovery window is a deletion with extra paperwork.

05

The uncomfortable finding nobody asks for

In most reviews, non-production is the biggest single recoverable line. Orphaned volumes, load balancers with no healthy target, GPU nodes held in case, and environments running 168 hours a week to serve about forty. We schedule them off, tag every survivor with an owner, and hand you the list of resources nobody was willing to claim. That conversation is awkward and it is where the money is.

34%

Year-one spend reduction

Median reduction in monthly cloud invoice across optimisation engagements, measured against the trailing quarter before review.

61%

Of savings from non-production

Share of recovered spend attributable to idle development, staging and abandoned experiment environments rather than production rightsizing.

0rebuilds

Applications rearchitected

Optimisation work delivered without asking engineering teams to refactor an application. Rewrites are a separate decision with a separate business case.

Advisory

Architecture review on retainer, before the purchase order.

A block of engineering hours for the decisions that are expensive to reverse: fabric topology, storage class, scheduler policy. Or whether the thing you are about to buy is the thing you need.

Design review before commitment

Bring the vendor's proposed topology and we will mark it up. Oversubscription ratios, failure domains, rebuild time at the proposed drive size, the single power feed nobody costed, and the cost of growing it by half. We are not bidding on that hardware, which is the point.

Workload profiling

We instrument the real job — an OpenFOAM case, a Nextflow pipeline, a vLLM serving path — and report where the wall-clock time goes. It is usually not where the procurement assumed. The finding often changes the shopping list more than the budget.

Operational readiness assessment

A written assessment of monitoring coverage, runbook completeness, restore evidence and patch currency — mapped to the Essential Eight where that is your obligation. You get a prioritised remediation list with effort estimates, whether or not you then engage us to do the work.

Second opinion on an estate we did not build

Retainer hours can be spent on someone else's architecture, including one that is currently failing. We will tell you what we would change, and whether the honest answer is a change of architecture or a change of expectation. Sometimes the design is sound and the deadline was never real.

Unit
Blocks of 40 engineering hours, drawn down against any of the four activities above.
Expiry
Unused hours roll forward one quarter and then lapse. We would rather you spent them than banked them.
Access
A named engineer, scheduled within five business days for non-urgent work.
Output
A written document every time, including the options we rejected and why. No verbal-only advice.

Incident Practice

How a P1 actually runs.

Written down in advance, because the middle of an incident is a poor time to invent a process. Every interval below is a design target for planning, not a contractual commitment.

Detection

An alarm fires with a runbook already attached, a synthetic transaction fails, or you telephone the NOC. All three paths open the same incident record with the same severity rules. There is no slower, second-class path for a fault a human noticed first.

Classification and acknowledgement

The analyst on shift classifies severity against a published matrix and acknowledges to a person, not to a ticket queue. P1 means the platform is unusable, losing data, or breaching a regulatory obligation. Nothing else earns it, and inflating severity to get attention is treated as a process defect.

Engage the right tier

Tier one holds the incident and executes the runbook. If the runbook does not resolve it within the first cycle, we escalate rather than retry the same step with more conviction. The design engineer at tier three joins any P1 by default — that is the commitment the rest of this page depends on.

Communications cadence

One incident channel, one named incident commander, and a written update every thirty minutes whether or not there is news. Silence is the most common complaint made about managed service providers, and it is entirely avoidable. An update saying we still do not know is a legitimate update.

Contain, then diagnose

Restoring service and finding the cause are separate jobs, done in that order. We will fail over, drain a node or roll a firmware level back to get you working. The diagnosis becomes harder as a result, and we accept that. Where evidence would be destroyed, we capture it first and say so in the log.

Restoration and verification

Closed means measured. The queue draining at its normal rate, the latency distribution back inside its envelope, failed jobs re-queued or formally accounted for. We verify against telemetry rather than against the absence of alarms, because an alarm can be silenced by the same fault that caused it.

Post-incident review, published

Within five business days you receive a written review: timeline, root cause, contributing factors, the changes we have made, and what we got wrong. If the cause was ours it says so in the first paragraph. We do not issue an unexpected combination of factors as a root cause, and we will not describe a capacity decision as an unforeseeable event.

On response times. Acknowledgement and engagement intervals quoted on this page are indicative planning figures. Contractual service levels, the point at which availability is measured, the exclusions and any service credits exist only in a signed services agreement. Ask us to put a number in that document if your business case relies on it.

Commercial Detail

Service levels, exit, and who can see your data.

The questions procurement asks after the technical evaluation is finished. Answered here so they are not a surprise in the contract negotiation.

Operations Handover

Send us your last three incident reviews. We'll tell you what your monitoring is missing.

No capability deck. We read the reviews, your alarm catalogue and the shape of your on-call roster. Then we write back with the gaps we can see, and which tier closes them. If the honest answer is that you do not need us yet, you will get that instead.

Existing clients — 24/7 NOC
+61 0 0000 0001
New engagements
hello@cloudnatives.example
Direct line
+61 0 0000 0000

Sydney · Melbourne · Canberra. Shifts worked in Australia, every hour of the year.