Restore customer service
Scheduler, checkpoint, failover, and maintenance systems recover the job. Zenobia does not sit in this critical path.
Minutes · handled by existing operationsFor infrastructure operators
Recover the customer job first. Then decide what the stranded machine can safely do next—without exporting the workload, risking the fleet, or starting a six-week research project.
Two clocks, deliberately separated
Zenobia begins after the existing fleet systems have moved the customer workload and isolated the suspect asset.
Scheduler, checkpoint, failover, and maintenance systems recover the job. Zenobia does not sit in this critical path.
Minutes · handled by existing operationsZenobia works on the drained machine until there is enough evidence to return, limit, repair, RMA, retire, or run one named test.
Measured by time to justified dispositionThe operator’s actual question
The deepest physical cause matters when it changes the safe operating state, the part to replace, fleet exposure, or entitlement to a supplier remedy.
The evidence supports normal service, with recurrence checks attached to the case.
Use the node only for the workloads and operating range it has shown it can run reliably.
The implicated part and supporting evidence justify removing the asset from service.
The decision is not yet safe; the exact next experiment and required capability are explicit.
How it fits into operations
Schedulers and recovery systems restore the customer job first.
Zenobia starts after removal from production, not in the critical service path.
Each result determines the safest high-value test to run next.
The operator gets an action, limits, evidence, and recurrence conditions.
If escalation is needed, the phenomenon and ruled-out alternatives are packaged.
The case checks whether the action worked and improves future triage.
What must be true
Incident capture has a fixed budget and never holds customer job recovery.
Confidential workloads are reduced to the smallest approved pattern that carries the phenomenon.
Only tests approved for the agreed operating envelope run; exceptional actions are separately authorized.
Raw workload data stays local. The operator releases the minimum evidence needed.
Every active test has hard bounds for power, temperature, timing, and other risk signals.
Recurrence monitoring checks whether the return, limit, repair, or RMA decision worked.
The operator’s economics
A successful deployment should move these operational measures, without increasing repeat failures.
Useful capacity returned rather than left stranded in diagnostic limbo.
Days from isolation to a justified disposition—not merely time to run a check.
Operator and vendor labor spent collecting, recreating, and explaining the incident.
Whether a returned node fails again under the same or related conditions.
Cases accepted without another cycle of missing logs, swaps, and reproduction requests.
Inventory and repair capacity tied up by machines that have not reached a decision.
Continue the investigation
Start with one hardware family and one existing quarantine workflow.