For infrastructure operators

Get the node
out of limbo.

Recover the customer job first. Then decide what the stranded machine can safely do next—without exporting the workload, risking the fleet, or starting a six-week research project.

Two clocks, deliberately separated

Service recovery should not wait for hardware closure.

Zenobia begins after the existing fleet systems have moved the customer workload and isolated the suspect asset.

Clock 01

Restore customer service

Scheduler, checkpoint, failover, and maintenance systems recover the job. Zenobia does not sit in this critical path.

Minutes · handled by existing operations
Node isolated
Clock 02

Close the hardware case

Zenobia works on the drained machine until there is enough evidence to return, limit, repair, RMA, retire, or run one named test.

Measured by time to justified disposition

The operator’s actual question

What should happen to this machine?

The deepest physical cause matters when it changes the safe operating state, the part to replace, fleet exposure, or entitlement to a supplier remedy.

Return

The evidence supports normal service, with recurrence checks attached to the case.

Return with limits

Use the node only for the workloads and operating range it has shown it can run reliably.

Repair, RMA, or retire

The implicated part and supporting evidence justify removing the asset from service.

Hold for one named test

The decision is not yet safe; the exact next experiment and required capability are explicit.

How it fits into operations

From isolated node to accountable action.

  1. 01

    Existing systems flag the node

    Schedulers and recovery systems restore the customer job first.

  2. 02

    The node is isolated

    Zenobia starts after removal from production, not in the critical service path.

  3. 03

    Approved experiments run

    Each result determines the safest high-value test to run next.

  4. 04

    A disposition is returned

    The operator gets an action, limits, evidence, and recurrence conditions.

  5. 05

    The support case is ready

    If escalation is needed, the phenomenon and ruled-out alternatives are packaged.

  6. 06

    The fleet learns

    The case checks whether the action worked and improves future triage.

What must be true

The investigation earns its place in the fleet.

Bounded production overhead

Incident capture has a fixed budget and never holds customer job recovery.

Synthetic reproducers

Confidential workloads are reduced to the smallest approved pattern that carries the phenomenon.

Warranty-aware execution

Only tests approved for the agreed operating envelope run; exceptional actions are separately authorized.

Operator-controlled disclosure

Raw workload data stays local. The operator releases the minimum evidence needed.

Automatic safety stops

Every active test has hard bounds for power, temperature, timing, and other risk signals.

Measured follow-through

Recurrence monitoring checks whether the return, limit, repair, or RMA decision worked.

The operator’s economics

Measure capacity recovered—not diagnoses produced.

A successful deployment should move these operational measures, without increasing repeat failures.

Recovered accelerator-hours

Useful capacity returned rather than left stranded in diagnostic limbo.

Time out of service

Days from isolation to a justified disposition—not merely time to run a check.

Human triage hours

Operator and vendor labor spent collecting, recreating, and explaining the incident.

Repeat incidents

Whether a returned node fails again under the same or related conditions.

First-pass RMA acceptance

Cases accepted without another cycle of missing logs, swaps, and reproduction requests.

Spares and repair load

Inventory and repair capacity tied up by machines that have not reached a decision.

Continue the investigation

Give the fleet a decision it can enforce.

Start with one hardware family and one existing quarantine workflow.