Put the machine in a justified state
Return it, limit it, repair or RMA it, or keep it isolated for one named test.
AI-factory yield management
Zenobia gets stranded hardware back into a justified state—with evidence for what the machine can safely do next.
It does not run the same checklist every time. Each result narrows the possible causes and tells Zenobia what to test next.
The whole product, in five levels
The customer sees a disposition. Underneath it is an investigation that persists, crosses organizational boundaries, and improves with every resolved case.
Return it, limit it, repair or RMA it, or keep it isolated for one named test.
Eliminate possible causes instead of repeating the same health-check sequence.
Hypotheses, experiments, exclusions, decisions, and evidence persist together.
The case can run in operator, vendor, or supplier environments under local control.
Closed cases improve approved tests, methods, and the next investigation.
The Zenobia agent
Zenobia keeps the original incident evidence, then perturbs, measures, and observes the isolated system thousands of times. The point is not more testing. The point is choosing the next useful test.
Oscillating calibrated sensors across the stack.
Per-rail, transient, phase-resolved response.
Register reads, ECC, retention, and stress patterns.
Latency, skew, and jitter across compute, network, and tray-control domains.

Synthetic, micro, and approved real-world workloads.
Hardware PMU, custom counters, and synchronized traces.
BIOS, BMC, drivers, policies, and governors.
Connections, lanes, links, switches, and dependencies.
Each measurement changes what Zenobia believes—and therefore what it does next.
From incident to disposition
A fixed suite asks the same questions in the same order. Zenobia spends the next experiment only where it can remove uncertainty.
Keep the device identity, workload behavior, and operating conditions from the original event.
Maintain several possible causes instead of jumping to the first plausible answer.
Run the safe experiment most likely to distinguish between the remaining causes.
Correlate the response across sensors, counters, registers, workloads, and matched controls.
Use the result to shrink the possibility set and decide which test matters next.
Produce a disposition, operating limits, evidence, and the next named test if uncertainty remains.
Stop condition: enough evidence to take a safe, commercially useful action—not an impressive pile of telemetry.
The case follows the hardware
The raw workload and proprietary tools stay where they belong. The case carries the smallest reproducible phenomenon, what has been ruled out, and what still needs to be tested.
Existing systems recover the job. The suspect node enters quarantine.
Zenobia works on the isolated machine within approved bounds.
Evidence, exclusions, and the reproducer move with the hardware.
Vendors and suppliers use their own protected tools and data.
A signed finding returns, the action is checked, and detection improves.
Engineering & outcome targets
Design-partner targets; not current guarantees.
Continue the investigation
Recover useful capacity, reduce repeat incidents, and give every party evidence it can act on.