An engineering desk with a hand lifting a grid-response recommendation sheet, an amber review row, server racks and a closed UPS cabinet.
Data CentersArtificial IntelligenceEngineeringReliability

A Data-Center Recommendation Needs a Withdrawal Rule

Ashutosh SinghalAshutosh SinghalAugust 8, 20266 min read

A data-center setting has two jobs that can pull in different directions: keep the facility connected through brief voltage disturbances, and transfer it to backup when a fault requires that response. A recommendation that solves the first problem has not yet earned the second. My design standard is to make both conditions explicit, including the evidence that can withdraw an apparently successful answer.

Our Veriprajna simulation puts that standard to work in a synthetic facility. Its equipment inventory, voltage events and outcomes are fixtures, not a customer installation or measured incident. The model-assisted run uses previously cached replies through a local bridge, rather than fresh inference. Within that limited example, an added proposed fault changes the final answer from a preliminary passing setting to no recommendation.

That reversal matters because the workflow has to preserve a failed acceptance test even when it already has a plausible setting and a fluent explanation for it.

Two reasons to transfer

The simulation models a fleet of UPS units, or uninterruptible power supplies. One transfer path counts qualifying voltage disturbances inside a rolling time window. Once enough strikes accumulate, a unit transfers to backup. Another path responds independently to a sufficiently deep and sustained voltage dip. Changing the counting setting leaves that second path in place.

The engineering tension is understandable without an equipment diagram. A sensitive counter can transfer during a succession of disturbances the facility is meant to ride through. A more tolerant counter can keep the modeled load connected, but it still needs to respond to the fault cases used to test protection. In this demo, acceptance requires both behaviors across a finite event library.

Those outcomes also need careful naming. Transferring a data center to backup removes its load from the utility grid; it does not by itself establish that the servers lost power. Conversely, retaining modeled grid load says nothing about whether a real battery has sufficient energy. A setting recommendation cannot borrow either conclusion from the other metric.

The initial search finds a candidate with five strikes in a 90-second window, using aggregate counting. It passes the base library. That gives the workflow a reason to continue testing the candidate, not a reason to stop questioning it.

A proposed fault falls between the paths

The cached model challenger adds a synthetic event labeled a progressive transformer-winding insulation fault. The useful evidence is its behavior inside the simulator, rather than the authority suggested by its name.

It contains four counted dips spaced more than 90 seconds apart. For the preliminary five-strike candidate, earlier strikes fall out of the window before enough can accumulate. Every dip also remains above the modeled 0.60 per-unit deep-sag threshold, where per unit means a fraction of nominal voltage. Neither transfer path catches the proposed event for that candidate.

The workflow then repeats the search over all 32 configured combinations of strike threshold, window and counting mode, using the four base faults plus the added proposal. No candidate satisfies both the benign-event and fault-event conditions. The final result is abstention, with no selected configuration. This is a withdrawal of the preliminary recommendation, not a measurement that every candidate failed every individual test.

Qualified simulation report showing no final configuration, zero of 32 acceptable configurations, and the failed proposed winding-fault scenario
The synthetic cached-reply run ends with no final configuration and 0 of 32 acceptable settings. The added fault is a simulator proposal; the report and raw JSON are not engineering certificates or verified equipment diagnoses.

An important design choice is visible here: the model can contribute a new case, but deterministic checks decide whether any candidate remains acceptable. A persuasive account of the proposed setting cannot restore the recommendation after those checks fail. The demo explainer provides the video and further context for this example.

Why not keep adjusting until something passes?

Widening a counting window sounds like a natural response to widely separated disturbances. Lowering a threshold sounds like another. Either changes which events cause a transfer, so either can also undermine the goal of riding through benign disturbances. A repair must be evaluated against both objectives, rather than judged solely by whether it catches the newly added event.

The demonstrated search already checks its finite menu and returns no answer. Expanding that menu would be a new experiment. It could identify another candidate, but success would still depend on the same acceptance conditions and on what the model represents. An exhausted menu is evidence about those searched choices, not proof that every possible equipment setting is unsuitable.

I prefer a visible unresolved result to selecting the least disappointing failure and presenting it as a configuration. That preference has a cost: a team receives no new setting from this run. It has to decide what evidence or modeling change would justify another search. The refusal earns its place by making that missing work explicit.

In an actual engineering process, withholding a proposed change also needs to be distinguished from operating existing equipment. This simulation issues no equipment commands. Its abstention does not establish that a facility should disconnect, that its current settings are safe, or that hardware must be replaced.

A counterexample needs scrutiny too

A difficult test can expose a gap in an acceptance procedure while still being a poor description of physical equipment. The name “winding insulation fault” does not establish a transformer diagnosis. This example shows a simulated event escaping two modeled transfer paths; it does not validate that event as a real electrical fault.

That leaves two separate decisions. Under the present test assumptions, the workflow has no acceptable recommendation. For a real facility, engineers would also need to assess whether those assumptions represent the equipment and disturbances that matter. Removing the challenge because it blocks an answer would conceal the first decision. Treating the challenge as physical proof would skip the second.

A useful next step is therefore to state what would resolve the uncertainty. If the proposed event is physically relevant, the model or the available settings may need revision before a recommendation can proceed. If it is not relevant, excluding it needs an engineering reason tied to the modeled scope. Either route should preserve the failed test and its disposition so that a later passing result can be understood.

Measured equipment behavior, battery energy and transfer timing remain outside this demonstration. A real setting change would need verified equipment settings, measured disturbance data, a suitable validated electrical model and independent engineering review. More fluent model output cannot supply those missing inputs.

Here is my short explanation of why I want the recommendation withdrawn when its supporting tests fail.

For an AI-assisted engineering workflow, I want the acceptance record to identify the current candidate, the tests it satisfies, the test that can remove it, and the uncertainty left after removal. A passing answer is useful only while its stated reasons continue to hold. When they no longer do, the workflow should preserve the objection and withdraw the recommendation before anyone treats it as permission to change equipment.

Related Research

Also Published On

Build Your AI with Confidence.

Partner with a team that has deep experience in building the next generation of enterprise AI. Let us help you design, build, and deploy an AI strategy you can trust.

Veriprajna Deep Tech Consultancy specializes in building safety-critical AI systems for healthcare, finance, and regulatory domains. Our architectures are validated against established protocols with comprehensive compliance documentation.